Description review
AI Evaluators: Assessing A Shopping Assistant
Terac · United States · back to the listing
HR standards
73/100
solid
Title ↔ description
54/100
needs work
Reads as
QA Engineer
85% confident
What this role officially is
software tester — ESCO, the EU occupation classification
Software testers perform software tests. They may also plan and design them. They may also debug and repair software although this mainly corresponds to designers and developers. They ensure that applications function properly before delivering them to internal and external clients.
Also known as: application software tester, unit tester, application tester, software application tester, tester, module tester
How others title the same work
Large employers
- Senior Software Quality Engineer Adobe
- Distributed Systems Testing Software Engineer, Python / Go Canonical
- Ubuntu Linux Kernel Test Engineer Canonical
- Distributed Systems Testing Software Engineer, Python / Go Canonical Ltd.
- Ubuntu Linux Kernel Test Engineer Canonical Ltd.
Startups
- Software Engineer, QA & Test Automation AviaryAI
What the listing never says
- No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency
The listing, marked up
Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.
What We're Researching
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
How It Works
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
Who This Is For
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
What You'll Do
• Review real user interaction traces with an AI shopping assistant
• Identify logical failures, inaccuracies, or poor recommendations in the text
• Create structured rubrics and verifiers to judge response quality
• Commit to 20+ hours per week of evaluation work on our internal platform
Who Should Apply
• Experience in data evaluation, quality assurance, or AI training
• Strong analytical skills with the ability to spot subtle errors in text
• Familiarity with e-commerce search and digital shopping experiences
• Ability to commit to a sustained workload of 20+ hours per week
Compensation
$50 per hour
Ready to participate?
Start your paid interview now
About Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at or on YouTube at @jointerac.
Originally posted on Himalayas
We're hiring AI evaluators to assess the accuracy and helpfulness of a new digital shopping assistant. This project focuses on understanding how well the system handles real-world e-commerce queries and where it falls short in its logic. Your analysis will directly feed into improving the underlying model and its response quality.
How It Works
You will review real interaction traces between users and the shopping assistant within our custom platform. As you analyze these conversations, you will pinpoint specific failures, logical errors, or unhelpful product recommendations. From there, you will create structured rubrics and verifiers to consistently judge future response quality. This is an ongoing remote engagement requiring 20+ hours per week.
Who This Is For
This opportunity is ideal for quality assurance specialists, AI data evaluators, and e-commerce professionals with a strong eye for detail. We welcome applicants with prior experience in prompt engineering, complex data annotation, or software testing. You should be comfortable analyzing text interactions deeply and building structured evaluation frameworks from scratch.
What You'll Do
• Review real user interaction traces with an AI shopping assistant
• Identify logical failures, inaccuracies, or poor recommendations in the text
• Create structured rubrics and verifiers to judge response quality
• Commit to 20+ hours per week of evaluation work on our internal platform
Who Should Apply
• Experience in data evaluation, quality assurance, or AI training
• Strong analytical skills with the ability to spot subtle errors in text
• Familiarity with e-commerce search and digital shopping experiences
• Ability to commit to a sustained workload of 20+ hours per week
Compensation
$50 per hour
Ready to participate?
Start your paid interview now
About Terac
Terac is building the world's largest pool of vetted human experts for AI. Researchers, AI labs, and product teams use Terac to recruit, screen, and pay study participants across industries, languages, and skill sets.
Learn more at or on YouTube at @jointerac.
Originally posted on Himalayas