hirly

Apply with hirly

Research Engineer, Frontier Data

Turing · Brazil

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

About Turing Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com . *This is a remote role and can be performed anywhere in Brazil/Colombia* The Role We are looking for a Research Engineer to help deliver frontier-quality datasets, RL environments, and evaluations that improve state-of-the-art models for a frontier AI lab client. You will work directly with the client's researchers and engineers, turning inbound requests and post-training goals into concrete technical proposals and data/environment specifications, and then owning the technical execution: building the quality and verification systems that ensure what we deliver meets extremely high standards for correctness, realism, diversity, difficulty, and measurable model lift. Most of this work is bespoke and custom, built for one client's evolving needs.You'll typically stay attached to a single frontier lab client so you can build real context and a working relationship with their team, with flexibility to support other accounts when needed. Because bespoke work is judged on both turnaround time and quality, both matter equally here. This role is designed for candidates with experience building and improving deep learning systems, especially where strong results depend on data quality, data curation, denoising, synthetic data generation, and rigorous evaluation. You'll operate in one or more of the following capability areas:

  • Coding and software engineering agents (repositories, unit tests, debugging, tool use, code reviews, long-horizon workflows)
  • RL environments and verifier-based training (tasks, rewards/verifiers, trajectories, evaluation harnesses)
  • Multimodal data and reasoning (text + images + documents + tables/charts; optional audio/video)
  • STEM reasoning (math, physics, chemistry, bio, engineering – solution verification and error analysis)
  • Modern embodied AI / VLM-driven agents (vision-language(-action) models, embodied task suites, tool/sensor/action abstractions, long-horizon interaction data) What You'll Do 1) Own data and environment quality from an AI researcher perspective
  • Respond to inbound requests from the client and translate ambiguous, evolving research goals into a technical proposal and clear data requirements: target skills, failure modes, difficulty calibration, coverage, and success metrics.
  • Provide the technical feasibility assessment, assumptions, acceptance criteria, and evaluation requirements needed to scope the contract or statement of work.
  • Define what “good” looks like by creating detailed rubrics, counterexamples, and boundary cases (what to include vs. exclude).
  • Perform deep, detail-oriented audits of produced data: spot subtle errors, reward hacking opportunities, leakage, ambiguity, inconsistent assumptions, and distribution shifts.
  • Drive iterative improvements using evidence: group failures (for example, by semantic similarity) to find the underlying gap, then turn that into concrete instructions for the production team — better few-shot examples, explicit counterexamples, clearer guidance on what's causing rejections. 2) Design and build datasets and RL environments for your capability area(s) Contribute to or lead the design of:
  • Task suites (single-step and long-horizon workflows)
  • Ground-truth signals (verifiers, unit tests, structured checks, reward functions, automatic validators)
  • Environment interfaces (APIs, tool schemas, state abstractions, database schemas, simulator-like dynamics) Depending on your mapped capability area(s), you may focus on:
  • Coding / SWE agents: data reflecting real development work (codebase navigation, bug localization, patching, tests, code reviews, CI-like constraints, refactors, security fixes).
  • Multimodality: tasks that test true multimodal reasoning (chart reading, document QA, UI understanding, diagram-based STEM reasoning, OCR-aware tasks).
  • STEM: tasks with verifiable solutions (symbolic checks, reference solvers, numerical validation, step consistency, unit sanity).
  • Modern embodied AI / VLM-driven agents: interaction data and environments for vision-language(-action) models. 3) Build robust validation, denoising, and synthetic data systems
  • Build project-specific verification systems that combine deterministic checks, model-based evaluators calibrated against gold sets, and targeted human review.
  • Implement automated validation and filtering to achieve frontier-grade signal-to-noise: deduplication, decontamination, leakage checks, consistency checks, difficulty and diversity controls.
  • Develop synthetic data generation and augmentation pipelines where appropriate: programmatic task generators, controlled perturbations, scenario templating, simulator-/tool-driven rollouts.
  • Create documentation and data cards: dataset intent, known limitations, recommended use, and evaluation linkage. 4) Use evaluations and training runs to prove impact
  • Design and run evals that reflect the client's intended usage.
  • Produce analysis that connects data to outcomes: pre/post comparisons, error breakdowns, ablations that identify which data attributes drive lift.
  • When needed, run in-house fine-tuning or RL-style experiments (or partner with research) to demonstrate that the data/environment improves model behavior in measurable ways. 5) Collaborate effectively with large production teams without being ops-heavy
  • Give technical direction to the distributed expert network producing the data, by providing clear specs, examples, edge cases, and fast feedback loops based on audits and quantitative signals.
  • You own technical translation, acceptance criteria, evaluator and verifier design, quality diagnosis, and the technical release recommendation. The Strategic Project Lead owns production execution, staffing, throughput, project cost, and recovery.
  • You are expected to be highly engaged in reviewing and improving outputs from large annotation/creation efforts, but not primarily responsible for hiring, staffing, or people operations — that sits with the SPL.
  • Communicate directly with the client on technical findings, quality risks, and recovery options. Align changes to scope, schedule, or external commitments with the account and delivery owners. Who We're Looking For
  • 4–5 years of experience building or improving deep learning systems where data quality mattered materially (training, post-training, evals, or agentic systems).
  • Strong intuition for the “data ingredients” that drive model improvements: what to collect, what to filter, what to synthesize, and how to measure.
  • Ability to communicate clearly with researchers and engineers: turning research objectives into concrete specs, and turning messy outputs into actionable insights.
  • Comfortable being the direct technical point of contact for a client, including explaining setbacks or trade-offs when things don't go as planned.
  • Demonstrated ability to be extremely detail-oriented in diagnosing subtle data quality issues and failure modes.
  • Solid programming ability with a bias for shipping: Python proficiency required; comfort with SQL/structured d