About the Role We're a fast-growing AI infrastructure company building the technical foundation for training and evaluating frontier AI agents. Our team includes International Olympiad medalists, serial AI startup founders, and researchers with publications at ICLR, NeurIPS, and similar top venues. We're hiring Research Engineers to work across quality control automation, benchmarks, and synthetic data — shaping how agents learn and improve at the frontier. This is a high-ownership, high-ambiguity role best suited to engineers who thrive without a prescribed roadmap and are energized by building systems that directly push the state of the art in RL-based agent training. What You'll Do Build systems for creating new environments, improving data quality, and translating real-world workflows into tasks and benchmarks. Design, run, evaluate, and iterate on agent training environments. Design experiments to understand model behavior, agent failure modes, and data quality issues. Develop tools that help researchers, engineers, and data vendors produce higher-quality tasks, trajectories, and feedback loops. Work across the full lifecycle of agent training data — from task design and environment setup to trajectory collection, evaluation, and validation. Partner with external vendors to identify bottlenecks and improve quality and throughput of the data engine. Build metrics and analyses to assess whether tasks, environments, and evals are genuinely useful for training frontier agents. What We're Looking For Required 2–4 years of relevant engineering experience. Strong proficiency in Python, Docker, and Linux environments. Experience with benchmarks and evals — able to reason about task realism, rubric reliability, environment usability, and trajectory quality for RL training. Sharp attention to detail; able to spot subtle inconsistencies in data, model behavior, or task design. Proven track record of building tools, pipelines, or research infrastructure with minimal guidance. Experience at an early-stage startup; comfort working independently in fast-paced, ambiguous environments. Experience designing metrics and validation workflows. Strong quantitative or technical foundation — demonstrated through competitive programming, published research, or substantial independent project work. Able to operate effectively in unstructured problem spaces and communicate clearly across time zones. Nice to Have Background in reinforcement learning, agent evaluation, or post-training data pipelines. Competitive programming achievements or olympiad-level problem-solving experience. Prior experience building or scaling data engines for ML/AI systems. Location This role is on-site in San Francisco, CA . Visa sponsorship is available. Compensation & Benefits Compensation details are competitive and commensurate with experience at an early-stage, well-funded AI company. Specific packages will be discussed during the interview process.
Full Stack Engineer ($140K–$200K + Equity) Building Data Systems for Scientific Research
CoffeeSpace
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic
Dev/Research Engineer
Cintal, Inc.
Research Engineer, Production Model Post-Training
Anthropic
Senior Research Engineer, Model Training / Pretraining
Intelix.AI
Lead Research Engineer, Data Quality
Hud