About Build AI Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world. Job Summary We’re hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You need to deeply understand the philosophy of evals: what an eval is allowed to claim, contamination, leakage, construct validity, and whether a number actually corresponds to a capability. More practically, you should have pushed consequential evals before — evals that changed what a lab trained, shipped, or collected, not a weekend leaderboard. Key Responsibilities Design and run benchmarks for video and world models, with Build data and without it, under the same protocol Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes Work with research and dataset/quality so collection and evals inform each other Push evals that are consequential enough that people change plans when the number moves Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries You may be a good fit if you have (Must-have qualifications) You have shipped or driven evals that mattered: they changed training, hiring, product, or data decisions You understand eval philosophy well enough to argue about validity, not only to plot a curve Ideally video, robotics, world models, or multimodal, but the bar is consequential evals more than a specific domain You will not confuse a pretty dashboard with an eval that is allowed to decide things Strong candidates may also have experience with (Nice-to-have qualifications) Video, robotics, world models, or multimodal evals You have designed evals used in a paper, a product launch, or a data decision Background in construct validity, contamination, or leakage Exposure to evaluation operations: throughput, failure taxonomy, repeatability Benefits Competitive pay Medical, dental, and vision packages with generous premium coverage $500 per month credit for waiving medical benefits Housing subsidy of $2k per month for those living within walking distance of the office Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan) Various wellness benefits covering fitness, mental health, and more Daily lunch and dinner in our office Unlimited compute budget subject to ROI justification Unlimited Codex and Claude credits Travel How we're different Build believes in the Bitter Lesson . By taking a general approach of learning from humans, our addressable market is all physical labor. We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed. Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: [email protected]
Member of Technical Staff, Post-Training Research
Intelligence
Member of Technical Staff, ML Engineer
Intelligence
Country Lead
Build Ai
Lead Engineer, Multimodal ML Data Infrastructure
Build Ai
Lead Scientist - Artificial Intelligence
Gevernova
Associate Director, Data Science
Merck