Backend Engineer — High-Volume Data Processing
- Moves you to
- United States
- Support
- Visa sponsorshipRelocation support
- Posted
- Oct 1, 2026
Backend Engineer — High-Volume Data Processing
Location: Dumbo, Brooklyn, NY
Work Arrangement: Onsite 5 days/week — non-negotiable
Employment Type: Full-time
Base Salary: 220K–300K
Equity: Competitive equity
Visa: Open to visa transfers (including OPT and H-1B transfers); additional sponsorship may be considered for the right candidate
Relocation: Up to $10K in relocation support
About the Opportunity
We're partnering with a fast-growing AI/data company that provides real, proprietary enterprise data to leading frontier AI labs.
The company acquires data generated through how businesses actually operate, transforms it into de-identified datasets that remain useful, and licenses those datasets for next-generation AI model development.
The business has scaled from $0 to a multi-eight-figure run rate in a matter of months and operates with a lean engineering team of approximately five. This is a high-impact, ground-floor opportunity to take significant ownership of critical systems at a rapidly scaling company.
They're looking for a Backend Engineer to own and evolve the critical data-processing pipeline that transforms sensitive enterprise data into AI-training-ready datasets.
This role is specifically about code applied to the data itself. It is not a traditional data engineering, ETL, or data-platform position focused primarily on moving data between systems.
What You'll Do
Build and scale multi-stage, high-volume data-processing pipelines that de-identify sensitive enterprise data from sources such as collaboration tools, email, cloud storage, codebases, and project-management systems
Own backend systems end-to-end — design, delivery, debugging, correctness verification, and safe recovery
Ensure silent failures, duplicated output, stale data, and other correctness issues are caught before customer delivery
Make retries, checkpoints, partial failures, replay, backfills, and rollbacks safe and understandable across asynchronous orchestration, batch workers, queues, and object storage
Build internal tooling and dashboards to understand pipeline performance, investigate data quality, and support release decisions
Collaborate closely with ML, full-stack, platform, and security teammates to process new data modalities and expand pipeline capabilities
What They're Looking For
Required
3–10 years of software engineering experience with personal ownership of production backend systems
Personally shipped and operated production backend systems, ideally multi-stage, high-volume data-processing pipelines
Relevant undergraduate STEM degree from a top-30 U.S. institution
- Strong distributed/asynchronous backend fundamentals, including experience with concepts such as:
Queues and workers
Concurrency
Idempotency
Retries
Partial failures
Recovery
Meaningful full-time experience at an early-stage or high-growth startup, operating with real ambiguity and ownership
Strong proficiency in a general-purpose backend language
Willing and able to work 5 days/week onsite in Dumbo, Brooklyn
The undergraduate education and onsite requirements are firm. Candidates who do not meet them will not be considered.
Technical Background
The current stack includes:
Python, AWS, Airflow, AWS Batch, SQS, S3, DynamoDB, PostgreSQL
Python experience is not required. Strong backend engineers coming from Go, Rust, or Java are also encouraged to apply.
Experience with AWS and an understanding of NLP and Named Entity Recognition (NER) are important for this work.
The team also values engineers who are creative around data — people who want to build tooling around pipelines, understand how systems are performing, investigate data quality and correctness, and solve the broader problem rather than simply ship backend code.
A backend-leaning full-stack mindset can be valuable.
Especially Relevant Experience
Strong candidates may also bring:
Experience at a highly regarded engineering organization paired with meaningful startup experience
Clear career progression and increasing ownership
AWS experience
Batch-processing or workflow-orchestration experience
Internal tooling, observability, or dashboard development around data-processing systems
Experience reasoning about correctness, recovery, and failure modes in production systems
Who This Role Is Not Designed For
This is unlikely to be the right fit for candidates whose experience is primarily:
Traditional data engineering, data platform, or infrastructure/platform work centered on moving data between systems
Databricks-style ETL without deep ownership of the processing logic itself
Frontend-heavy full-stack development without significant backend depth
Long-term big-company engineering without meaningful early-stage/high-growth startup experience, unless the work is a direct domain match
Startup exposure limited primarily to internships or brief engagements
Team & Culture
This is a small, high-conviction team that moves quickly and holds a high bar for ownership and quality.
The culture emphasizes:
Understanding the fundamental problem and solving it
Following through on commitments
Bias toward action and learning quickly
Producing world-class work
Humility and supporting teammates
Focusing on outcomes rather than outputs
With a small engineering team and significant customer demand, this is an environment for someone who wants substantial ownership rather than a narrowly defined piece of a mature system.
Compensation & Benefits
Base Salary: 220K–300K
Equity: Competitive equity
Medical / Dental / Vision: 100% covered
PTO: Unlimited
Relocation Support: Up to $10K
Visa Sponsorship
The company is open to visa transfers, including OPT and H-1B transfers.
Additional sponsorship may be considered for the right candidate. Sponsorship is not guaranteed and will be evaluated based on the individual candidate.
Client Interview Process (after candidates are submitted to the client)
Agency/recruiter screening and the client's initial resume review are not counted as interview stages.
Stage 1 — Intro Call | 30 minutes
Conversation with the hiring manager covering your background, motivation, and alignment with the opportunity.
Stage 2 — System Design | 1 hour
Distributed-systems/data-processing exercise focused on architecture, tradeoffs, and correctness in high-volume backend pipelines.
Stage 3 — Take-Home Technical Screen | A Few Hours
Practical backend exercise focused on fundamentals and hands-on coding. Completed independently and reviewed by the engineering team.
Stage 4 — Final Round Onsite | 2 hours
Onsite interviews at the Dumbo, Brooklyn office with engineering and leadership, focused on final technical and team evaluation and giving you an opportunity to get to know the company and team.
Candidates who are not currently in New York will be flown in and provided accommodation for the final onsite.