This one is closed
Live roles like this one
-
B
4h ago
Freelance Mechanical CFD Engineer - AI Trainer
Mindrift Saudi Arabia $37/hr
- C 6d ago
-
D
6d ago
OpenSanctions Berlin, Germany / REMOTE (CET ±3
- C 1w ago
See every "Python Engineer Freelance" role →
Get new “Python Engineer Freelance” roles by email
One email a day with what is new in "Python Engineer Freelance". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade This listing scored 39/100, which is an F. It lost the most ground on pay transparency. See the breakdown
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Pay transparency 12 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Corroboration 5 / 10 Whether more than one source carries this listing.
- Freshness 4 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
- Role specificity 0 / 10 Whether the listing is tagged well enough to tell what the role actually is.
-10 Ghost-job penalty — Deducted for signals that this posting may not be a real, currently-open role — staleness, repeated relisting, or talent-pool language.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves
Frontier coding agents are already good at passing tests. We measure whether they pass them the right way.We're building a dataset to evaluate the safety and conduct of AI coding agents — not just whether they complete a task, but whether they do it honestly, within scope, and without cutting corners.You'll design tasks where the easy path is the unsafe one, and write the tests that catch it:
- Build realistic developer environments — a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks that pair a benign development goal with a tempting unsafe shortcut: scope creep, policy violations, data corruption, over-permissive changes
- Write tests that verify whether the agent completed the task the right way — catching corners cut, not just checking outputs
- Iterate on tasks and tests based on QA feedback — review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT:
- Not data labeling;
- Not prompt engineering;
- Not cybersecurity or red-teaming — there is no attacker in the scenario. Cybersecurity experience is a nice-to-have but not a requirement. We're looking for engineers who understand how code should behave, not penetration testers. Strong software engineers, not security specialists;
- Not writing code from scratch — the agent writes most of the code; you design the situation and evaluate the outcome;
What we look for
- 4–5+ years in software development;
- Core stack: Python, JavaScript/TypeScript;
- Strong test design skills — functional and integration tests that separate safe from unsafe completion, not just correct from incorrect;
- Hands-on experience with coding agents (Claude Code, GitHub Copilot CLI, Codex, or similar);
- Familiarity with GitHub PRs and CI workflows as a user;
- Stack breadth is welcome, not a filter. Tasks simulate real repositories with databases, CI pipelines, and deploy scripts, so wider backend and infrastructure exposure is genuinely useful — but you don't need to be an expert in every layer;
- English proficiency — B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. The real difficulty is building the temptation — a scenario where the unsafe or out-of-scope path is the path of least resistance — and then writing tests that reliably catch an agent that took it. Tasks have many valid solutions; tests must accept all of them and reject the bad ones.
How it works
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Project time expectations
For this project, tasks are estimated to require around 20-25 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.
Compensation
On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.
Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.
Originally posted on Himalayas
Apply for this role Opens himalayas.app — the link as listed; we have not yet verified it is the employer's own page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 03 Aug 2026 Himalayas first sighting
Seen on 1 board over 0 days.