Lead AI Engineer, Model Training & Evaluation
- Hiring from
- Singapore
- Work type
- Remote
- Posted
- Sep 24, 2026
Lead AI Engineer Chipforge
Remote (Singapore/Australia-based) · Full-time
About us
Chipforge is agentic AI for chip design. We build autonomous, closed-loop design and verification agents that turn RTL generation, verification, and design automation into a compounding loop, one that gets faster and sharper with every project. At the centre is our purpose-built model family that iterate alongside engineers rather than simply assisting them.
Chipforge spans FPGA through ASIC delivery for fabless and design-services teams who need to move faster without losing control of their IP. We are the agentic layer that makes every engineer who touches them faster, from first RTL draft through to verification.
About the role
Chipforge is hiring a Lead AI Engineer to take ownership of our purpose-built model family for RTL generation, verification, and design automation. This is a hands-on role: you will drive model training, fine-tuning, and evaluation day to day, and work directly alongside our existing model engineering team to accelerate its development. You will report into Chipforge's engineering leadership.
▍WHAT YOU'LL DO
- Own the day-to-day training and fine-tuning of our model family, including post-training with rewards drawn from simulation, lint and synthesis results, alongside preference-based methods where they fit
- Run the loop from evaluation to improvement: analyse where our models fail, target those weaknesses with verified synthetic data, and measure whether each change actually helped
- Set the evaluation standard our models are measured against, extending our existing benchmarks for RTL generation and verification
- Own the policy for training data quality and contamination control, setting the standards our data pipeline must meet and making sure benchmark results stay clean and credible
- Train our models to use EDA tools and design workflows reliably, working with the agent team on the interfaces those models call
- Run the training and inference infrastructure our models use, making pragmatic build-versus-fine-tune-versus-buy calls appropriate to a lean, resource-constrained team
- Work with Chipforge's hardware and EDA domain experts to encode verification and design knowledge into training data and evaluation criteria
- Convert research ideas into production-ready features on a realistic delivery timeline, partnering with customer-facing teams to turn field feedback into priorities
- Work closely with our existing model engineering team, setting technical direction across training and evaluation
- Keep across relevant LLM and agentic-AI research and bring anything genuinely useful into our model roadmap
▍WHAT YOU'LL NEED
- Hands-on experience fine-tuning and evaluating large language models for a specific technical domain, not only general-purpose chat or assistant use cases
- Experience building training data for LLMs, including generating, verifying and filtering synthetic data at scale
- Direct experience with post-training, ideally reinforcement learning from automated feedback (for example GRPO or RL on code execution results); experience with RLHF or DPO also counts
- Strong Python and PyTorch (or JAX) fundamentals, with real experience running training and inference workloads at meaningful scale
- Demonstrated ability to build domain-specific evaluations and diagnose failures, going beyond aggregate scores to find out why a model fails and what to change
- A track record of shipping production-grade ML systems, not only research prototypes
▍WHAT SETS STRONG CANDIDATES APART
We're a lean team, so this role needs someone who can own training and evaluation end to end. You'll stand out if you have:
- Designed, run and rigorously analysed your own experiments
- Built and maintained production systems, beyond research notebooks and one-off scripts
- A working understanding of transformer and language model internals, enough to diagnose why a training run behaves the way it does
- Scaled training or inference across multiple GPUs or distributed setups
- Built your own tooling to make the team faster
- A track record of taking responsibility for the downstream impact of the systems they build
- Designed reward signals from automated checks such as test results or compiler output, and an understanding of how models learn to game them
▍NICE TO HAVE
- Experience with agentic architectures, tool-use interfaces, or multi-agent workflows
- Familiarity with RTL design and verification methodologies (UVM, formal verification), synthesis, place-and-route, or PPA optimisation
- Exposure to EDA tools such as Cadence or Synopsys, or open-source equivalents
- Understanding of FPGA-plus-ASIC development workflows
- Prior technical leadership or mentoring experience, ideally at a startup
- Publications or contributions at venues such as NeurIPS, ICLR, ICML, DAC, ICCAD, or DVCon
- Experience partnering with customer-facing or product teams in a deployment-driven (not purely research) environment
▍HOW WE WORK
This is a remote-first role, open to candidates based anywhere in Singapore or Australia. We work closely as a distributed team across Australia, Singapore, and India, so comfort collaborating across time zones matters. Occasional travel for team or product sessions may come up, but the day-to-day is remote.