Posted
07 Jul 2026
Last seen
07 Jul 2026
Location
Brazil
Lifecycle
mature
Grade
F
Data Science at TRACTIANThe Data Science team at TRACTIAN focuses on extracting valuable insights from vast amounts of industrial data. Using advanced statistical methods, algorithms, and data visualization techniques, this team transforms raw data into actionable intelligence that drives decision-making across engineering, product development, and operational strategies. The team constantly works on optimizing prediction models, identifying trends, and providing data-driven solutions that directly enhance the company’s operational efficiency and the quality of its products.What you'll doWe're looking for a Data Engineer with a strong engineering foundation and comfort with AI workflows to join our Data Foundry team. In this role, you'll be the bridge between our model training and data annotation teams, building the pipelines and infrastructure that turn raw, messy data into gold-standard datasets ready for AI consumption.ResponsibilitiesDesign and maintain robust data pipelines to ingest from a wide range of sources, including APIs, documents, websites, and raw sensor dataIntegrate and optimize ETL/ELT processes developed by MLE colleagues, improving performance, reliability, and long-term maintainabilityOwn the full dataset lifecycle, from raw ingestion through cleaning, validation, and delivery as training-ready dataDefine and enforce data quality standards and governance practices across the Data Foundry teamBuild and maintain labeling pipeline infrastructure for ML applications, working closely with the annotation teamParticipate in architectural decisions, code reviews, and technical mentorship within the teamDocument data sources, pipeline logic, and processing decisions for reproducibility and team alignmentRequirements3+ years of experience in data engineeringDegree in Computer Science, Data Engineering, Computer Engineering, Information Systems, or equivalent technical backgroundSolid understanding of the ML training lifecycle and what properties make a dataset suitable for model trainingFamiliarity with layered data architecture patterns such as Medallion Architecture (Bronze/Silver/Gold) or Data MeshProficiency in Python, with focus on data manipulation, pipeline development, and automationWorkflow orchestration using code-based tools such as Temporal, Airflow, Prefect, Dagster, or equivalentDistributed data processing with Spark, Databricks, or similarREST and gRPC API integrationStrong SQL skills, both for data modeling and query optimizationExperience with streaming systems and event-driven pipelines (Kafka, Kinesis, or equivalent)Soft SkillsComfortable jumping into ongoing codebases and optimizing work built by others, without needing to start from scratchTechnology-agnostic: you evaluate tools based on what the project needs, adopt new ones quickly, and don't get attached to a specific stackAt ease in fast-moving environments where priorities shift and the right answer isn't always obviousEngineering-first mindset: you think in pipelines, own outcomes, and care about the quality of what you shipDriven by curiosity and innovation, not by comfort with a known toolsetNice to HaveExperience making architectural decisions and contributing to the technical growth of a team, formally or informallyGo, for high-performance pipeline componentsdbt for transformation layer modelingOpen table formats: Delta Lake, Apache Iceberg, or HudiData quality frameworks such as Great Expectations or SodaCloud experience, preferably OCI (our current migration target). AWS, GCP, or Azure background is also valuedRapid prototyping with Streamlit or similar tools. The use of LLMs and GenAI to speed up internal tooling and experimentation is actively encouragedExperience with data annotation workflows or training dataset pipelinesOriginally posted on Himalayas