Backend Engineer – High-Volume Data Processing_SO
- Moves you to
- United States
- Support
- Visa sponsorshipRelocation support
- Posted
Show job descriptionHide job description
We are looking for a Backend Engineer to own and evolve the critical data-processing pipeline at a fast-growing AI data company. You will build the systems that turn sensitive enterprise data into de-identified, AI-training-ready datasets for leading frontier AI labs. This role is about code applied to the data itself: it is not a traditional data engineering, ETL or data-platform position focused on moving data between systems. You will own backend systems end to end, from design and delivery to debugging, correctness checks and safe recovery, making sure problems are caught before data reaches customers. As one of roughly five engineers, you will have real ground-floor ownership of core systems at a company that is scaling very quickly. It is a great fit for a backend engineer with strong distributed-systems fundamentals, startup experience and a genuine curiosity about data quality and correctness.
Details:
- Schedule: Full time
- Location: Onsite, 5 days/week – Dumbo, Brooklyn, New York (non-negotiable)
- Start: ASAP
- Duration: Long-term
- English: Fluent
- Type of collaboration: Internal Employment
- Work authorization: Open to visa transfers (including OPT and H-1B transfers); new sponsorship may be considered for an exceptional candidate but is not guaranteed
- Relocation: Up to $10K relocation support
- Benefits: Competitive equity, 100% covered medical/dental/vision, unlimited PTO
About the project:
The client is a fast-growing AI data company that supplies real, proprietary enterprise data to leading frontier AI labs. It acquires data generated by how businesses actually operate – from collaboration tools, email, cloud storage, codebases and project-management systems. That data is transformed into de-identified datasets that stay useful and are then licensed for next-generation AI model development. The business went from $0 to a multi-eight-figure run rate in a matter of months. It runs with a lean engineering team of about five people, so each engineer owns large parts of critical systems rather than a narrow piece of a mature platform. The culture values solving the real problem, following through on commitments, moving fast, humility and focusing on outcomes rather than outputs.
You have:
- 3–10 years of software engineering experience with personal ownership of production backend systems
- Hands-on experience shipping and operating production backend systems, ideally multi-stage, high-volume data-processing pipelines
- An undergraduate STEM degree from a top-30 U.S. university (firm requirement)
- Strong distributed/asynchronous backend fundamentals: queues and workers, concurrency, idempotency, retries, partial failures and recovery
- Meaningful full-time experience at an early-stage or high-growth startup, working with real ambiguity and ownership (not just internships or short stints)
- Strong skills in a general-purpose backend language – Python is a plus, but strong Go, Rust or Java engineers are welcome
- AWS experience and an understanding of NLP and Named Entity Recognition (NER)
- A creative approach to data: you like building tooling around pipelines, digging into data quality and correctness, and solving the wider problem
- Willingness and ability to work onsite 5 days/week in Dumbo, Brooklyn (firm requirement)
- Nice to have: experience at a highly regarded engineering organization combined with startup experience; clear career progression; batch-processing or workflow-orchestration experience (e.g. Airflow, AWS Batch); internal tooling, observability or dashboards for data systems; a backend-leaning full-stack mindset
- Tech stack: Python, AWS, Airflow, AWS Batch, SQS, S3, DynamoDB, PostgreSQL
This role is not a fit if your experience is mainly:
- Traditional data engineering, data-platform or infrastructure work centered on moving data between systems
- Databricks-style ETL without deep ownership of the processing logic
- Frontend-heavy full-stack development without strong backend depth
- Long-term big-company engineering without meaningful early-stage or high-growth startup experience
What to do:
- Build and scale multi-stage, high-volume pipelines that de-identify sensitive enterprise data from collaboration tools, email, cloud storage, codebases and project-management systems
- Own backend systems end to end: design, delivery, debugging, correctness checks and safe recovery
- Catch silent failures, duplicated output, stale data and other correctness issues before customer delivery
- Make retries, checkpoints, partial failures, replays, backfills and rollbacks safe and easy to understand across async orchestration, batch workers, queues and object storage
- Build internal tools and dashboards to track pipeline performance, investigate data quality and support release decisions
- Work closely with ML, full-stack, platform and security teammates to handle new data types and expand pipeline capabilities
Interview process:
1. Intro call – 30 minutes with the hiring manager about your background and motivation
2. System design – 1 hour, a distributed-systems/data-processing exercise on architecture, trade-offs and correctness
3. Take-home technical screen – a few hours of hands-on backend coding, reviewed by the engineering team
4. Final onsite – 2 hours at the Brooklyn office with engineering and leadership (travel and accommodation covered for candidates outside New York)