About the Role We're a fast-growing AI infrastructure company building the technical foundation for training and evaluating frontier AI agents. Our team includes International Olympiad medalists, serial AI startup founders, and researchers with publications at top venues (ICLR, NeurIPS, and similar). We're looking for Research Engineers to work across agent quality control automation, benchmarks, and synthetic data — shaping how AI agents learn and improve. This is a high-ownership, high-impact role at an early-stage company. You'll work in ambiguous, fast-moving problem spaces alongside a tight-knit team where your contributions directly influence the trajectory of frontier AI development. What You'll Do Build systems for creating new environments, improving data quality, and translating real-world workflows into tasks and benchmarks. Build systems for creating, running, evaluating, and improving agent training environments. Design experiments to understand model behavior, agent failure modes, and data quality issues. Develop tools that help researchers, engineers, and data vendors create higher-quality tasks, trajectories, and feedback loops. Work across the full lifecycle of agent training data — from task design and environment setup to trajectory collection, evaluation, and validation. Partner with external vendors to identify bottlenecks and improve the quality and throughput of the data engine. Build metrics and analyses to assess whether tasks, environments, and evals are genuinely useful for training frontier agents. What We're Looking For Required: 2–4 years of relevant engineering experience. Proficiency in Python, Docker, and Linux environments. Experience with benchmarks and evals, including reasoning about task realism, rubric reliability, environment usability, and trajectory quality for RL training. Strong attention to detail — ability to spot subtle inconsistencies in data, model behavior, or task design. Track record of building tools, pipelines, or research infrastructure with minimal guidance. Early-stage startup experience; comfort working independently in fast-paced, ambiguous settings. Experience designing metrics and validation workflows. Strong quantitative or technical foundation, demonstrated through competitive programming, research, or independent project work. Ability to thrive in unstructured problem spaces and communicate clearly across time zones. Nice to Have: Background in reinforcement learning or AI alignment research. Experience working with large-scale data pipelines or vendor ecosystems. Publications or contributions to open-source ML/AI tooling. Compensation & Benefits Salary: $150,000 – $250,000 USD annually, depending on experience. Visa sponsorship is available. Equity participation in an early-stage, well-resourced AI company. Location This is an on-site role based in San Francisco, CA . Candidates should be prepared to work in-person with the team.
Research Engineer, Synthetic Data
Clera
Forward Deployed Research Engineer
Clera
Senior Machine Learning Research Engineer
Carnaby Fox
Full Stack Engineer ($140K–$200K + Equity) Building Data Systems for Scientific Research
CoffeeSpace
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic
Dev/Research Engineer
Cintal, Inc.