Lead Data Engineer
- Hiring from
- Australia
- Work type
- Hybrid
- Posted
530,570 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
About Maincode
Maincode makes Matilda, an Australian-made AI stack we own end to end, from tin to token. We own and operate our own infrastructure here in Australia, we train and maintain our own frontier models, and we run multiple surfaces on top of them, including chat, coding and API. Everything we ship is built on capability we control ourselves. We are hiring someone who cares about that. You should be passionate about Australian-made technology, a builder first, and keen to execute at the frontier of data engineering and model training.
The role
A frontier model is only as capable as the data it learns from, and every model carries the worldview of the data it was trained on. We did this for Australia with Matilda’s voice. Now we are building the capability to do it for any country or market: training open weights models on custom datasets that reflect the cultural, regulatory, governance, and ideological requirements of the people who will use them.
You will be the hands on engineer who builds the system that makes those datasets. You will own the pipeline end to end, from web scale crawling of hard to reach, country specific sources through to filtering, deduplication, quality control, and the agent driven pipelines that turn raw material into training sets. We need someone who moves fast, writes code every day, and knows how to use AI agents to multiply what a small team can produce without letting quality slip.
This is an engineering role, not a management one. You will set the technical direction, and as the work grows you will help us hire the engineers who join you.
What you will do
Build, run, and scale crawlers and scrapers for web scale collection, including specialized crawlers for high value, hard to reach sources such as legislation, regulator guidance, government publications, local media, and community forums.
Develop robust pipelines for extraction, filtering, cleaning, deduplication, and formatting to ensure the highest quality inputs for model training.
Build agentic pipelines where AI agents generate, annotate, critique, and verify data at scale, with the evaluation and spot checks needed to trust what they produce.
Turn a market’s requirements, from its laws and regulators to its cultural norms and governance expectations, into dataset specifications, data mixtures, and evaluations we can measure against.
Work directly with our core training team to iterate on data mixtures, evaluate data quality, and respond to model performance feedback.
Track provenance, licensing, privacy, and security for every dataset you ship.
What you will bring
A track record of building, scaling, and running large scale web crawling or scraping systems in production.
Strong engineering in Python and at least one of Go, Rust, or C++, with real experience in distributed data processing (Spark, Ray, Beam, or similar) and petabyte scale storage.
Hands on knowledge of the hard parts of crawling: JavaScript heavy sites, headless browsers, rate limiting, robots.txt, and access and licensing constraints.
Experience using LLMs and AI agents inside data pipelines for classification, extraction, synthetic data generation, or quality scoring, and a clear sense of where they fail.
A bias for shipping. You get a working pipeline running in days, then make it reliable.
Experience leading the technical direction of a project and bringing other engineers along with you.
Highly regarded
Direct experience building pre-training or post-training datasets for large language models, including SFT, preference, or reinforcement learning data.
Experience with multilingual, regional, or culturally specific data, or with evaluating models against local norms and regulation.
Experience fine-tuning or post-training open weights models.
Open source contributions to crawling, data processing, or dataset tooling.
A strong network within the Australian technology, academic, or media ecosystem to facilitate unique data partnerships.
Eligibility
Must have the right to work in Australia.
Willingness to undergo background checks as required.
Details
Location: Melbourne, VIC (or Hybrid/On-Site across Australia).
Type: Full time, permanent.
Remuneration: Negotiable based on experience, plus super and equity.
Who thrives here
Engineers who understand that high quality data is the true moat of any frontier model, and that it is also where a model’s values come from. If you care about Australian-made technology and you would rather spend your week shipping crawlers, tuning agent pipelines, and reading samples of your own data than sitting in steering committee meetings, you will do well here.