Description review

Member of Technical Staff (Data Intelligence)

Reka · Singapore, UK, USA · back to the listing

HR standards

59/100

needs work

Title ↔ description

70/100

solid

Reads as

Unclear

no confident match

What this role officially is

data scientist — ESCO, the EU occupation classification

Data scientists find and interpret rich data sources, manage large amounts of data, merge data sources, ensure consistency of data-sets, and create visualisations to aid in understanding data. They build mathematical models using data, present and communicate data insights and findings to specialists and scientists in their team and if required, to a non-expert audience, and recommend ways to apply the data.

Also known as: data scientists, data engineer, research data scientist, data expert, data research scientist

How others title the same work

Large employers

  • Applied AI Engineer Automattic
  • Machine Learning Engineer, CX Intelligence Coinbase
  • AI Engineer - FDE (Forward Deployed Engineer) Databricks
  • AI Engineer - FDE (Forward Deployed Engineer) - U.S. Federal Sector Databricks
  • Senior AI Engineer – Notebooks Datadog

Startups

  • Attendi | Medior Machine Learning Engineer | Amsterdam, Netherlands | ONSITE (hybrid) | €6,000 - €7,000 per month | Full-time (80–100%, ~4–5 days/week) | Visa sponsorship + 30% ruling possible Attendi
  • Member of Technical Staff (applied) Anthrogen
  • Aptura AI | Full-Time | MTS (Applied AI), MTS (SWE / Product) | London | ONSITE / HYBRID Aptura AI
  • Arcforma AI (arcforma.ai) | AI Engineer (Marketing / Construction / Arcforma AI (arcforma.ai)
  • Staff AI Engineer - Agent Architecture & Behavior Artisan

What the listing never says

  • 21 bullet points. Long requirement lists deter qualified candidates, who read them as hard gates. Scope clarity
  • No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency

The listing, marked up

Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.

In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get.

What you’ll do

•
Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds

•
Explore open source datasets and create internal ones most suitable to build fundamental World Models

•
Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data.

•
Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs

•
Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction

•
Track and optimize throughput, storage, and compute utilization across pipelines and related assets

What we’re looking for

•
Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems

•
Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems

•
Demonstrated research experience with data compositions, quality, and dataset releases

•
Ability to design and execute experiments with convincing unbiased outcomes

•
Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)

•
Solid Python skills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)

•
Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale

•
Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers

•
Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments

Reka's Mission

Reka's mission is to build useful multimodal artificial intelligence and use it to empower organisations and businesses. We are a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world. Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI over the past decade.

Why Reka?

•
An Elite Team: Collaborate with top-tier engineers, researchers, operators from renowned organizations like Google DeepMind and Facebook AI Research (FAIR) and successful startups, driving innovation in cutting-edge AI technology.

•
Massive Market Opportunity: Be part of a rapidly growing industry poised to transform multiple sectors globally, offering the chance to make a significant impact.

•
Mission-Driven Environment: Work alongside a collaborative, mission-focused team dedicated to advancing AI for meaningful applications.

•
Inclusive and Open Culture: Thrive in an open and inclusive work environment that values diverse perspectives and fosters creativity.

•
Generous Benefits: Enjoy 5 weeks of paid leave to recharge, comprehensive healthcare benefits including vision and dental, and additional perks that support your well-being.

•
Visa Support: We provide visa assistance, including H1B and OPT transfers, for US employees to ensure a smooth transition and support your career with us.

How this was produced

Highlights are found by rule, not by a model: each one is a phrase matched at a known position, and every note is a template we wrote. The two scores come from a typed-decision model (Jev) that reads the listing against the official role definition and real listings for the same role, and returns probabilities rather than prose — it never writes any of the words on this page, and never chooses what to highlight.

Deterministic penalty applied to the HR score: 8 points (from 67 before penalties). Reviewed 21 Sep 2026.