About the Role Join a lean, high-caliber engineering team building a synthetic data pipeline that turns domain-specific workflows into scalable training tasks for AI agents. You'll work alongside Olympiad medalists and published researchers, with direct ownership over generation methods, validation systems, and quality metrics that expand what AI models can do. What You'll Do Design and build end-to-end synthetic data pipelines that transform domain-specific workflows into structured, realistic, and challenging training tasks. Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains. Develop synthetic task generation methods that produce diverse, realistic, and learnable outputs at scale. Build automated tooling to mutate, validate, and continuously improve synthetic task quality. Analyze model and agent performance on synthetic tasks to understand what they teach and where they break down. Define and implement metrics to quantify task diversity, realism, learnability, and overall data quality. What We're Looking For 2–4 years of experience in software engineering, ML engineering, or AI research with a focus on data pipelines, ML infrastructure, or synthetic data systems. Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications. Proficiency in Python and comfortable working in Linux environments with containerization tools such as Docker. Strong understanding of synthetic data quality criteria — diversity, realism, learnability — and an honest awareness of its limitations. Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or environments for AI agents or large language models. Track record of independently owning and delivering technical projects end-to-end with minimal predefined requirements. Sharp eye for detecting edge cases, inconsistencies, or quality issues in algorithmically generated datasets. Ability to reason from first principles about task design, scoring, and failure modes. Comfortable thriving in unstructured, early-stage environments where the roadmap isn't fully written yet. Familiarity with reinforcement learning, agentic AI workflows, or LLM post-training pipelines is a plus. Compensation & Benefits Salary range: $150,000 – $250,000 USD annually. Visa sponsorship is available. Location On-site in San Francisco, CA, USA . Singapore candidates are also considered.
Research Engineer, Synthetic Data
hud
Research Engineer, Synthetic Data
HUD
Member of Technical Staff, Synthetic Data
TypeSafe AI
Opal Electronics — iOS Engineer
Davidjoseph Co
Principal Software Engineer - Back End (Wildfire)
Paloaltonetworks
Principal Software Engineer (Data Platform Prisma AIRS)
Paloaltonetworks