Research Engineer, Synthetic Data
- Salary
- $150K–$250KUSD per year
- Moves you to
- Singapore
- Support
- Visa sponsorship
- Posted
- Sep 27, 2026
About the Role
This is a Research Engineer role focused on synthetic data, sitting within a roughly 15-person engineering team of Olympiad medalists and published researchers. You will build the pipelines that turn domain-specific workflows into scalable, high-quality training tasks for AI agents, directly shaping what models learn and how well they perform.
What You'll Do
Build end-to-end synthetic data pipelines that transform domain-specific workflows into realistic, structured, and challenging training tasks.
Collaborate with subject-matter experts to create synthetic tasks for AI agents across professional and technical domains.
Design task generation methods that produce diverse, realistic, and learnable outputs at scale.
Build tooling to mutate, validate, and continuously improve synthetic tasks.
Analyze model and agent performance on synthetic tasks to identify what they teach and where they break down.
Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.
What We're Looking For
2 to 4 years of experience in software engineering, machine learning engineering, or AI research, with hands-on work building data pipelines, ML infrastructure, or synthetic data systems.
Proficiency in Python and experience developing in Linux environments using containerization tools such as Docker.
Demonstrated experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
Strong understanding of synthetic data quality criteria and evaluation metrics, including diversity, realism, and learnability, as well as their inherent limitations.
Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
Track record of independently owning and delivering technical projects end-to-end with minimal predefined requirements.
Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.
Sharp eye for edge cases, inconsistencies, and quality issues in synthetic or algorithmically generated data.
Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines is a plus.
Comfortable operating in unstructured, early-stage environments and reasoning from first principles.
Strong communication skills for asynchronous, cross-timezone collaboration.
Compensation & Benefits
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
Location
On-site in Singapore.