About the Role We're a fast-growing AI infrastructure company building the technical foundation for training and evaluating frontier AI agents. Our team includes International Olympiad medalists, serial AI startup founders, and researchers with publications at top venues including ICLR and NeurIPS. As a Research Engineer , you will work across agent training environments, benchmarks, and synthetic data pipelines — shaping how AI agents learn, improve, and are evaluated at scale. This is an on-site role based in San Francisco, CA (with a presence also in Singapore). Visa sponsorship is available. What You'll Do Build systems for creating new environments, improving data quality, and translating real-world workflows into tasks and benchmarks. Design, run, evaluate, and iteratively improve agent training environments. Design experiments to understand model behavior, agent failure modes, and data quality issues. Develop internal tools that help researchers, engineers, and data vendors produce higher-quality tasks, trajectories, and feedback loops. Work across the full lifecycle of agent training data — from task design and environment setup through trajectory collection, evaluation, and validation. Partner with external vendors to identify bottlenecks and improve the quality and throughput of the data pipeline. Build metrics and analyses to assess whether tasks, environments, and evaluations are genuinely useful for training frontier agents. What We're Looking For Required: 2–4 years of professional experience as a Research Engineer or in a comparable role delivering systems for AI agent training and evaluation. Strong proficiency in Python, Docker, and Linux environments. Hands-on experience with benchmarks and evaluations for RL training, including reasoning about task realism, rubric reliability, environment usability, and trajectory quality. Experience building systems for creating, running, and evaluating AI agent training environments. Experience designing experiments and building metrics/analyses to diagnose model behavior, agent failure modes, and data quality. Proven ability to build infrastructure, tools, or pipelines in early-stage environments without fully prescribed roadmaps. Nice to Have: Experience building internal research infrastructure or data pipelines. Experience designing metrics and validation workflows. Background in competitive programming, Olympiad mathematics/computing, academic research, or exceptionally strong independent project work. Comfort working and communicating clearly across time zones. Compensation & Benefits Salary: $150,000 – $250,000 USD annually, depending on experience. Equity participation in a well-funded, early-stage company (Series A/B stage). Visa sponsorship available. Opportunity to work alongside a world-class team at the frontier of AI agent research and infrastructure. Location Primary: San Francisco, CA, United States — on-site. Singapore office also available for candidates based in Southeast Asia. This is not a remote role.
Research Engineer, QC Automation
Clera
Research Engineer, Benchmarks
Clera
Dev/Research Engineer
Cintal, Inc.
Senior Research Associate, Bioengineering
ODDITY LABS
Concrete Research Engineer
Genex Systems
Hydraulics Research Engineer
Genex Systems