Senior Data Engineer
- Salary
- $180K–$220KUSD
- Hiring from
- United States
- Work type
- Remote
- Posted
- Sep 29, 2026
Senior Data Engineer - Legal tech
Location: Remote US, strong preference for New York City
Salary: $180,000-$220,000 | Equity: 0.15%-0.25%
Visa transfers accepted
Our client is building a high-performance agentic litigation engine that combines millions of embedded legal documents, proprietary judicial behavioural intelligence, and specialised AI agents to transform how litigation strategy is developed.
You would be joining a highly technical, product-focused team working at the intersection of legal data, artificial intelligence, and large-scale data infrastructure.
The Role
Our client is looking for a Data Engineer with 4+ years of experience to take ownership of the caselaw and docket data layer. You will build production-grade pipelines that ingest, normalise, and enrich data from PACER, NYSCEF, state court systems, and published opinions.
The legal data environment is complex and highly unstructured. You will work with millions of court records, PDFs, scanned filings, XML, HTML, and other semi-structured sources, turning raw information into reliable, queryable data that powers the company's AI strategy layer.
This is a high-impact role where every downstream product surface depends on the quality and reliability of the systems you build.
What you'll own
- Build and maintain production data pipelines ingesting millions of court records from PACER, NYSCEF, state court systems, and published opinions.
- Take full ownership of the caselaw and docket data layer.
- Convert PDFs, scanned filings, XML, HTML, and other semi-structured inputs into clean, queryable data structures.
- Design LLM-assisted extraction workflows using Claude, Gemini, or OpenAI.
- Turn unstructured legal text into reliable structured data and actionable signal.
- Work closely with full-stack engineers to support judicial behavioural intelligence and AI strategy products.
- Run statistical analyses and optimise SQL queries to improve data quality and performance.
- Build and improve RAG pipelines and semantic search capabilities using tools such as Pinecone and Voyage AI.
- Operate and optimise the AWS data stack, including S3 and RDS.
- Develop data infrastructure in line with SOC 2 subservice architecture requirements.
What we're looking for
- 4-8 years of experience building and operating production data pipelines.
- Strong Python and SQL skills.
- Experience working with PostgreSQL, MongoDB, or similar databases.
- Experience processing large-scale, messy, or semi-structured datasets.
- Familiarity with PDF parsing, XML and HTML scraping, and document extraction workflows.
- Experience designing or operating LLM-powered data extraction pipelines.
- Understanding of NLP, RAG pipelines, semantic search, and vector databases.
- Experience with AWS services, particularly S3 and RDS.
- Familiarity with dbt and modern data engineering practices.
- Strong analytical and problem-solving skills.
- Ability to work closely with software engineers and translate complex data into reliable product infrastructure.
- Previous experience working with legal, court, caselaw, docket, or other highly regulated data would be highly valuable.
- Ability to maintain at least four hours of daily overlap with Eastern Time.
Bonus points
- Experience working directly with PACER, NYSCEF, state court systems, or published legal opinions.
- Experience building data products for legal technology, litigation, compliance, or regulated industries.
- Experience with Pinecone, Voyage AI, Claude, Gemini, or OpenAI APIs.
- Experience operating data systems within SOC 2 environments.
Benefits
- Salary of $180,000-$220,000.
- 0.15%-0.25% equity.
- Remote US-based working.
- Visa transfers supported, including OPT and H1B transfers.
- Opportunity to own a critical data layer powering an AI litigation platform.
- Direct impact across legal data infrastructure, AI products, and judicial intelligence.
If you are interested, send me a message or apply directly.