About the Role This role sits at the core of a small, technical team building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You'll own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand true agent capability. Your work directly shapes the credibility and rigor of a product at the frontier of AI evaluation. What You'll Do Design, implement, and own the quality of internal benchmarks for evaluating frontier AI agents on domain-specific tasks. Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks. Build reliable infrastructure to run models and agents against benchmark tasks at scale. Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes. Validate that benchmark performance correlates meaningfully with real-world agent behavior and customer needs. Write clear documentation and benchmark reports that make results legible and credible to technical audiences. What We're Looking For 2–4 years of experience in research engineering or ML engineering, with a focus on building and delivering AI benchmarks, evaluation infrastructure, or agent environments. Strong proficiency in Python, Docker, and Linux environments for research or production infrastructure. Experience designing and running evaluations for AI agents or large language models. Experience building and operating infrastructure to reliably run AI models or agents against benchmark tasks at scale. Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation. Ability to collaborate with subject-matter experts and translate domain workflows into evaluation criteria. Strong attention to detail and a habit of spotting subtle inconsistencies and edge cases. Comfort working independently in fast-paced, early-stage startup environments with unstructured problem spaces. Excellent written communication skills for cross-functional collaboration across time zones. Bonus: Published papers or technical writing on AI benchmarking, model evaluation, or failure modes; experience with RL training pipelines or widely used public benchmark projects. Compensation & Benefits Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available. Location On-site in San Francisco, CA, United States .
Research Engineer, Benchmarks
hud
Research Engineer, Benchmarks
HUD
Opal Electronics — iOS Engineer
Davidjoseph Co
Principal Software Engineer - Back End (Wildfire)
Paloaltonetworks
Principal Software Engineer (Data Platform Prisma AIRS)
Paloaltonetworks
Manufacturing Engineer
Schreiberfoods