Senior Rust Engineer - AI Code Evaluation (Codex / Claude Code, up to $200/hr)
G2i Canada, United States
This one is closed
Live roles like this one
-
B
17h ago
Senior Technical Services Analyst
First Due United States $110k/yr
-
C
3w ago
Pre-Sales Engineer, North America
QuestDB New York, New York, US
- C 1d ago
-
B
1d ago
Senior Software Engineer (Python)
Full Fibre Limited United Kingdom £75k - £90k/yr
See every "Rust Engineer AI" role →
Get new “Rust Engineer AI” roles by email
One email a day with what is new in "Rust Engineer AI". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade This listing scored 39/100, which is an F. It lost the most ground on pay transparency. See the breakdown
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Pay transparency 12 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Corroboration 5 / 10 Whether more than one source carries this listing.
- Freshness 4 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
- Role specificity 0 / 10 Whether the listing is tagged well enough to tell what the role actually is.
-10 Ghost-job penalty — Deducted for signals that this posting may not be a real, currently-open role — staleness, repeated relisting, or talent-pool language.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
Senior AI Interaction Evaluator (Codex / Claude Code)
Contract | $100–$200/hour | 10–20 hrs/week | Start ASAP (through early May)
Check out this Loom video for more details!
We’re looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
This is not a traditional engineering role.
You won’t be writing production code.
You’ll be evaluating something harder: whether the model thinks like a great engineer.
What This Role Actually Is
You will assess how AI coding agents behave in real-world scenarios — focusing on:
Whether the response makes sense
Whether the preamble and reasoning are useful
Whether the output reflects strong engineering judgment
Whether the interaction feels right to an experienced developer
This role is about engineering taste — not syntax correctness.
What You’ll Be Doing
Evaluate AI-generated coding interactions end-to-end
-
Judge whether outputs are:
Useful
Correct (at a high level)
Aligned with how a strong engineer would think
Assess the quality of explanations and reasoning, not just code
Distinguish between different levels of response quality (e.g. what makes something a 2 vs 4)
-
Provide clear, opinionated feedback on:
What worked
What didn’t
What felt “off” or misleading
Help define what great looks like when interacting with tools like Cursor
What We Mean by “Taste”
We’re specifically looking for engineers who can answer questions like:
Does this feel like something a strong engineer would actually say?
Is this explanation helpful, or just technically correct?
Is the model guiding the user well, or just dumping output?
Would this interaction build or erode trust?
You should be comfortable making subjective but rigorous judgments.
Who You Are
Staff / Principal-level engineer (or equivalent experience)
-
Strong background in one of the below:
TypeScript / JavaScript
Python
-
Hands-on experience using:
OpenAI Codex
Claude Code
Cursor
Deep familiarity with modern AI-assisted dev workflows
Able to evaluate code without needing to fully execute or deeply review every line
Comfortable giving direct, opinionated feedback
High bar for what “good engineering” looks like
Nice to Have
Experience with tools like Cursor or similar AI-first IDEs
Prior exposure to prompt design or evaluation workflows
Experience mentoring senior engineers or defining engineering standards
Engagement Details
Rate: $100–$200/hour
Hours: ~10–20 hours/week
Duration: Through early May (with possible extension)
Start: ASAP
-
Process:
Take-home evaluation exercise
One behavioral interview
Originally posted on Himalayas
Apply for this role Opens himalayas.app — the link as listed; we have not yet verified it is the employer's own page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 15 Aug 2026 Himalayas first sighting
Seen on 1 board over 0 days.