Description review

AI Evaluation Specialist

micro1 · Australia, Canada, Ireland, New Zealand, United Kingdom, United States · back to the listing

HR standards

64/100

needs work

Title ↔ description

62/100

needs work

Reads as

Unclear

no confident match

What this role officially is

software tester — ESCO, the EU occupation classification

Software testers perform software tests. They may also plan and design them. They may also debug and repair software although this mainly corresponds to designers and developers. They ensure that applications function properly before delivering them to internal and external clients.

Also known as: application software tester, unit tester, application tester, software application tester, tester, module tester

How others title the same work

Large employers

  • Senior Software Quality Engineer Adobe
  • Distributed Systems Testing Software Engineer, Python / Go Canonical
  • Ubuntu Linux Kernel Test Engineer Canonical
  • Distributed Systems Testing Software Engineer, Python / Go Canonical Ltd.
  • Ubuntu Linux Kernel Test Engineer Canonical Ltd.

Startups

  • Software Engineer, QA & Test Automation AviaryAI

What the listing never says

  • No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency

The listing, marked up

Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.

Role Title: AI Evaluation Specialist

Role Type: Contractor

Location: Remote (US, CA, UK, IE, AU, NZ)

micro1 is engaging AI Evaluation Specialists to assess and elevate the quality of AI assistant outputs for an enterprise AI training initiative. In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world input. No prior experience in AI is required — your domain knowledge is what matters.

Scope of Work

• Evaluate AI-generated outputs against detailed rubrics and defined quality standards, focusing on accuracy, relevance, and adherence to guidelines.

• Apply consistent, impartial judgment across a high volume of examples, ensuring a fair and reliable assessment process.

• Identify reasoning gaps, tool-use failures, or logic errors in AI assistant responses, providing actionable feedback for iterative improvement.

• Produce clear, concise written feedback on both strengths and areas for improvement, directly influencing model refinement and AI adoption practices.

• Participate in discussions regarding rubric interpretation and evolving quality standards, contributing to process optimization and best practices.

• Maintain meticulous documentation of evaluations and recommendations, ensuring transparency and traceability in assessment workflows.

Preferred Qualifications

• Experience in grading, quality assurance, editorial review, assessment, annotation, or similar fields demanding careful analysis and detailed feedback.

• Advanced, daily use of AI assistants (such as ChatGPT, Claude, or similar) as an essential work and productivity tool.

• Demonstrated ability to synthesize complex information and communicate findings effectively in writing.

• Background in process improvement, rubric development, or operational quality assessment in an enterprise or educational context.

• Strong critical thinking skills with a focus on consistency, integrity, and fairness in evaluations.

• Comfort working independently on large volumes of similar examples while maintaining high attention to detail.

• Collaborative mindset for sharing insights, discussing ambiguous cases, and refining evaluation criteria as models evolve.

Originally posted on Himalayas

How this was produced

Highlights are found by rule, not by a model: each one is a phrase matched at a known position, and every note is a template we wrote. The two scores come from a typed-decision model (Jev) that reads the listing against the official role definition and real listings for the same role, and returns probabilities rather than prose — it never writes any of the words on this page, and never chooses what to highlight.

Deterministic penalty applied to the HR score: 4 points (from 68 before penalties). Reviewed 21 Sep 2026.