About Build AI Build AI is the data hyperscaler for Physical AI. We're vertically integrated across hardware, manufacturing, logistics, collection, and model training to scale the physical labor dataset orders of magnitude faster than anyone in the world. Job Summary We’re hiring a lead for the data platform: camera on a worker to training-ready datasets, and out to research customers. Collection is monocular 1920×1080p 30fps in the wild, targeting 100M hours. The roles we mean: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, Eventual/Daft. This is not a warehouse, analytics, or generic backend seat. Key Responsibilities Own the data platform end-to-end: on-device capture, upload under flaky bandwidth, object storage, training-ready shards Compression, codecs, and storage-tier trade-offs so 1080p30 hours stay cheap enough to keep collecting Upload that survives bad networks: on-device buffering, batching, retries, a drop rate you can actually see Object storage and training-shard formats. The hard problem is petabyte-scale media, not a warehouse Own dataset packaging, versioning, and delivery to external research customers Work with Shenzhen firmware so new devices speak one ingest contract, not a custom path per SKU Make health, cost, and drop rate obvious as we add sites and countries You may be a good fit if you have (Must-have qualifications) You have owned a production media or sensor data path at real scale: object storage at petabyte scale, video codecs and compression, upload under flaky bandwidth, or training-shard / dataset formats That kind of data path: Tesla Autopilot, Waymo, Cruise, Zoox, Nuro, Samsara, Verkada, Netflix encoding, YouTube ingest, Scale, or Eventual/Daft. Demo-scale ETL is not this job Strong software engineering. Python and at least one systems language. Linux You measure cost and throughput, not whether the demo uploaded You want to scale in-the-wild physical-labor video, not run a generic data org Strong candidates may also have experience with (Nice-to-have qualifications) Pose, multi-camera, or other large media besides video Cloud (AWS or GCP), orchestration (Kubernetes, Airflow, Temporal), or IaC Dataset management or annotation tooling You have shipped dataset delivery to external research or training customers Benefits Competitive pay Medical, dental, and vision packages with generous premium coverage $500 per month credit for waiving medical benefits Housing subsidy of $2k per month for those living within walking distance of the office Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan) Various wellness benefits covering fitness, mental health, and more Daily lunch and dinner in our office Unlimited compute budget subject to ROI justification Unlimited Codex and Claude credits Travel How we're different Build believes in the Bitter Lesson . By taking a general approach of learning from humans, our addressable market is all physical labor. We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed. Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai
Lead Data Engineer – Physical AI Platform, Data Engineering
Cat
Senior Data Engineer – Physical AI Platform, Data Engineering
Cat
Senior Data Engineer
Boeing
Senior ML Accelerator Engineer - GPU
Generalmotors
Information Systems, IT, Data Science Intern - Summer 2027
Aerospace
Information Systems, IT, Data Science Intern - Summer 2027 (U.S. Person Required)
Aerospace