Data Engineer (Python, SQL, LLM Extraction) | Court and Caselaw Data | Remote US | Up to $220K + equity
- Hiring from
- United States
- Work type
- Hybrid
- Posted
- Sep 29, 2026
Mike Ross had a photographic memory. You'll need something rarer.
On television, the paperwork always turns up in the right folder, perfectly legible, exactly when the hero needs it. In real life, the paperwork is a scanned PDF from a county courthouse, skewed by three degrees, stamped twice and annotated in someone's handwriting.
Somebody has to turn that page into clean, reliable, queryable data. We think that person might be you.
We're recruiting for an early-stage legal technology company building AI products for litigators. Everything those products do rests on the court data layer underneath: what goes in, how cleanly it's structured, and how far it can be trusted. It's one of the most important hires the company is making right now, and they're ready to move quickly for the right person.
The job, in plain English
You will own the caselaw and docket data layer, start to finish. No handing your work to someone else's roadmap; this one is yours.
- Build production pipelines pulling from federal and state court systems, e-filing platforms and published opinions, at the scale of millions of records
- Turn PDFs, scans and uneven XML or HTML into clean, structured, queryable data
- Design LLM-assisted extraction workflows and, just as importantly, prove they're accurate
- Work with full-stack engineers so your data powers the product's analytical and AI features
- Run statistical analysis, tune SQL, and use RAG and semantic search to get more from the data
- Look after a cloud data stack within enterprise security and compliance guardrails
Your three big technical skills
- Python and SQL for production pipelines. Your pipelines run unattended, recover gracefully, and don't wake anyone at 3am. Your SQL holds up under scrutiny.
- Document extraction. PDF parsing, OCR and semi-structured text. Court data rarely arrives as neat rows, and you know it.
- Applied LLM work. RAG, semantic search and LLM-assisted extraction, with a healthy scepticism about model output. Vector database experience is a bonus.
The legal and data experience that gets you noticed
Direct experience with legal or court data is strongly preferred, and it saves months of ramp. Dockets, court records, opinions and federal filing systems all count, as does any work with regulatory or government data.
You don't need a law degree. You do need scar tissue. You know the same judge can appear under a dozen spellings, that case captions change mid-case, and that a field which looks reliable often isn't. On the data side, you've spent your career with large, dirty datasets and you're quick at cleaning them and finding the signal.
What this role is
- Real ownership of the layer the whole product depends on
- Hands-on, build-and-ship engineering in a small team
- A place where domain depth is rewarded and messy data is the interesting part
- Fully autonomous work: you set your own direction
What this role is not
- Not BI or analytics engineering on a tidy, pre-built warehouse
- Not a prompt-writing or research role; LLMs are a tool inside your pipeline, not the job
- Not for anyone who treats LLM extraction as plug-and-play, or who can't review the code an AI assistant writes
- Not a junior role: 4 to 8 years of production pipeline experience
- Not a managed role with a backlog handed to you each morning
- Not a law degree role: legal qualifications aren't expected
- Not open to candidates outside the US, or without Eastern Time overlap
The practical bits
- Location: Remote within the US, East Coast preferred. At least four hours of daily overlap with Eastern Time, and quarterly in-person onsites.
- Compensation: Up to $220K base plus meaningful equity
- Visas: Open to visa transfers
- Team: More than one hire planned
- Education: Bachelor's in computer science, data science or a related field
- Process: Three steps, including a live SQL and pipeline design session and a technical deep dive on pipelines you've built
Still reading?
If a badly scanned filing strikes you as a puzzle rather than a problem, you're exactly who we're looking for. Send us your CV and tell us about the messiest dataset you've ever tamed.
The title no longer says "East Coast preferred", because option 1 had dropped it. It's still in the practical bits, so candidates see it before they apply.