This one is closed
Live roles like this one
-
B
2h ago
Senior Data Engineer, Data Management & BI
NBCUniversal United States $115k - $145k/yr
-
C
6h ago
D2L Canada
-
B
4h ago
Freelance Mechanical CFD Engineer - AI Trainer
Mindrift Saudi Arabia $37/hr
-
B
6h ago
Senior Data Engineer (Snowflake)
Keyrus Portugal, Spain €45k - €50k/yr
See every "AI Evaluation Engineer" role →
Get new “AI Evaluation Engineer” roles by email
One email a day with what is new in "AI Evaluation Engineer". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade This listing scored 44/100, which is a D. It lost the most ground on pay transparency. See the breakdown
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Pay transparency 12 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Corroboration 10 / 10 Whether more than one source carries this listing.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Freshness 4 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
- Role specificity 0 / 10 Whether the listing is tagged well enough to tell what the role actually is.
-10 Ghost-job penalty — Deducted for signals that this posting may not be a real, currently-open role — staleness, repeated relisting, or talent-pool language.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
- Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT
- Not data labeling
- Not prompt engineering
- Not writing code from scratch - the agent writes most of the code; you guide and evaluate
What we look for
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
How it works
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Compensation
Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.
Originally posted on Himalayas
Apply for this role Opens himalayas.app — the link as listed; we have not yet verified it is the employer's own page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 07 Aug 2026 Himalayas first sighting
Seen on 1 board over 14 days.