Data Research Engineer
Chipforge.aiAbout Chipforge
Chipforge takes engineers from plain-English design intent to synthesizable, verified RTL, starting with FPGAs and expanding to ASICs, cutting out the time, iteration cost, and tool complexity that make that path slow today. We run in environments ranging from shared cloud to fully air-gapped deployments for customers with strict security and data-residency requirements.
About the role
This role owns the data behind our model work. You'll decide what goes into a training set and what doesn't, build the pipelines that produce it, and design the checks that establish whether a training example is any good in a domain where correctness is expensive to establish.
We're looking for someone who has built training data before, has views on how to tell good from bad, and can turn those views into something that runs automatically instead of relying on a human reviewer.
You'll work alongside the engineer who runs our training, the hardware engineers who are the authority on whether a design is correct, and the team that owns evaluation. You're not expected to be the hardware expert but you'll need familiarity with digital design and Verilog/VHDL. You are expected to get what's in their heads into the pipeline.
What you'll own
- The training corpus end to end: sourcing, licence filtering, deduplication, preparation, and versioned datasets our training runs can reproduce.
- Synthetic data generation with verification in the loop. Producing training examples, checking them against real tooling rather than a model's opinion, and keeping only what survives.
- The quality bar itself. Designing, calibrating, and maintaining the automated filters that decide what enters a training set, and catching drift before it reaches a training run.
- Data sources beyond public code, including signal generated by our own tooling, with the privacy and data-residency boundaries that come with customer environments.
- Contamination control and held-out construction, so our evaluation results mean what we say they mean.
- Translating hardware judgment into code. Getting the correctness criteria out of the engineers who hold them, encoding them, and applying them at a scale no reviewer can reach.
- The data roadmap, working directly with our Head of Product Engineering on how it lines up with what the product needs to ship.
What you'll need
- You've built synthetic training data with automated verification in the loop, not just prompt-and-collect. Generation, checking, filtering, and a defensible answer to how you knew the output was good.
- Judgment about data quality, and the ability to explain it. We'll ask you to look at real examples and tell us which ones you'd throw away and why.
- Strong production Python and pipeline engineering. Object storage, dataset versioning, orchestration, and jobs that produce the same result twice.
- Practical understanding of how training data reaches a model, and of what has to change when the target model does.
- Deduplication and contamination control, including near-duplicate detection and held-out overlap.
- Working knowledge of open-source licence compliance for training data.
- Comfortable reading code in a language you don't write. Much of our corpus is Verilog and SystemVerilog. You won't be expected to write RTL, and hardware engineers are available to you, but you need to read it well enough to tell a broken sample from a good one and to ask them a precise question.
- Working knowledge of LLMs: how they behave, where they fail, and how to build reliable systems around a component that isn't deterministic.
- Comfort operating with ambiguity. You'll be making calls that don't have a fixed playbook yet, and defending them.
Nice to have
- Reinforcement learning from verifiable reward, rejection sampling, or preference data collection.
- Hands-on continued pretraining or fine-tuning experience, including distributed training data formats.
- Exposure to hardware design workflows or EDA tooling, such as Verilog simulation or synthesis. You don't need a hardware background, but curiosity about the domain helps.
- Evaluation and benchmarking experience, particularly around code generation.
- Experience building for air-gapped or on-premises deployment environments.
- Prior experience mentoring other engineers. This role is likely to grow in that direction.
How we work
This is a remote-first role, open to candidates based anywhere in Australia. We work closely as a distributed team across Australia, Singapore, and India, so comfort collaborating across time zones matters. Occasional travel for team or product sessions may come up, but the day-to-day is remote