This one is closed
Live roles like this one
-
Binance Singapore, Taiwan 02 Oct 2026 fresh
-
eHealth, Inc. United States $130k - $162k/yr 01 Oct 2026 fresh
-
Sprinto India 01 Oct 2026 fresh
-
Nagarro Mexico 01 Oct 2026 fresh
See every "Software Engineer AI" role →
Get new “Software Engineer AI” roles by email
One email a day with what is new in "Software Engineer AI". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade
This listing scored 79/100, which is a B. It lost the most ground on remote clarity.
- Pay transparency 25 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Freshness 15 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Role specificity 6 / 10 Whether the listing is tagged well enough to tell what the role actually is.
- Corroboration 5 / 10 Whether more than one source carries this listing.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
Senior Contractor
Contract | Remote, worldwide | $100-$200/hour depending on experience and location | 10-20 hrs/week
Check out this Loom video for more details:
We're looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
This is not a traditional engineering role. You won't be writing production code. You'll be evaluating something harder: whether the model thinks like a great engineer.
What this role actually is
You will assess how AI coding agents behave in real-world scenarios, focusing on:
• Whether the response makes sense
- Whether the preamble and reasoning are useful
- Whether the output reflects strong engineering judgment
- Whether the interaction feels right to an experienced developer
This role is about engineering taste. Syntax correctness is the easy part.
What you'll be doing
- Evaluate AI-generated coding interactions end to end
- Judge whether outputs are useful, correct at a high level, and aligned with how a strong engineer would think
- Assess the quality of explanations and reasoning, not just the code
- Distinguish between levels of response quality (what makes something a 2 vs a 4)
- Give clear, opinionated feedback in writing on what worked, what didn't, and what felt off or misleading
- Help define what great looks like when working with tools like Cursor, Codex and Claude Code
What we mean by "taste"
We're looking for engineers who can answer questions like:
- Does this feel like something a strong engineer would actually say?
- Is this explanation helpful, or just technically correct?
- Is the model guiding the user well, or just dumping output?
- Would this interaction build or erode trust?
You should be comfortable making subjective but rigorous judgments, and explaining them clearly.
Who you are
- Senior, Staff or Principal-level engineer (or equivalent experience)
- Strong background in TypeScript/JavaScript or Python
- Hands-on experience with at least one of OpenAI Codex, Claude Code or Cursor
- Deep familiarity with modern AI-assisted dev workflows
- Able to evaluate code without executing it or reviewing every line
- Strong written and spoken English (B2 or above). You'll be writing detailed feedback and recording short video explanations, so this is a hard requirement
- Comfortable giving direct, opinionated feedback
- High bar for what good engineering looks like
Nice to have
- Prior exposure to prompt design or evaluation workflows
- Experience mentoring senior engineers or defining engineering standards
Engagement details
- Rate: $100-$200/hour depending on experience and location
• Hours: 10-20 hours/week, flexible
- Duration: ongoing. Projects run from about two weeks to a few months each, and we offer new ones to evaluators who do well
- Start: as soon as you clear the take-home and a project has an open seat
- Process: one take-home evaluation exercise with a recorded Loom walkthrough. No interview
Originally posted on Himalayas
Apply for this role Opens himalayas.app — the link as listed; we have not yet verified it is the employer's own page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 30 Sep 2026 Himalayas first sighting
Seen on 1 board over 0 days.