Substrate Bio logo

Member of Technical Staff, Intelligence

Hiring from
United Kingdom
Work type
Hybrid
Posted
Is this job info correct?

525,677 remote jobs, straight from company career pages

100% free · New jobs every hour

Show job description

LOCATION: King’s Cross, London · PATTERN: Hybrid, with regular time in the lab

The opportunity

We’re a stealth startup in AI and bio, building infrastructure for data generation. The team is small and elite, the problems are hard, and the foundations are being laid right now. You’d help write them from the first line, with real ownership and a direct line to the founders.

This role builds our analysis from the start. Every result leaves an instrument as a trace or a vendor file, and it has to become a value someone can rely on, linked back to the sequence and sample it came from. That means fitting it properly, checking it against controls, catching the plate or reagent lot that drifted, and showing the evidence. Someone has to own that work. That is this role.

About us

AI for biology has a data problem. The data that matters most doesn’t exist yet, so we’re building the infrastructure to produce it, with quality and traceability built in from day one.

We’re in stealth and heads-down on execution. We’ll share more about what we’re building once you’ve spoken with the team.

The role

You will join the intelligence team and own the analysis of our assay data, and the data and ML infrastructure it runs on, from the point the data has been captured to the result that gets delivered. You will work most closely with the protein characterisation scientists who run the assays, and with the software team, who get data out of the instruments and own the platform it lives on. You build on top of that platform, and you own the quality of the analysis that runs on it.

The role is part data engineer and part biostatistician, and all of the data is biological. Much of it is building the analysis and ML pipelines that run on our captured data, and the statistics inside them. The rest is presenting those results clearly, with the statistics that matter, such as uncertainty and QC, shown in a form they can understand. The value is in doing that carefully, in code that behaves the same way on every run.

What you will own

● Analysis and ML pipelines that turn captured assay and sequence data into consistent, versioned datasets, linked to the sequence, sample, plate and well each result came from, built with the software team who own data capture.

● Data validation before any analysis runs, catching a mislabelled plate, a swapped sample or a missing well.

● Sequence-level bioinformatics, such as computing the properties of each protein sequence and checking constructs for problems before the DNA is ordered.

● Biostatistics for experimental data: experimental design, replicates and controls, error propagation, outlier handling, and assay acceptance criteria, such as CV across replicates, Z′ for plate assays, and reference standards staying within range.

● Curve fitting, including dose-response (EC50 and IC50) with sound normalisation and fits to SPR, BLI, HPLC and nanoDSF traces, with a measure of confidence on each fitted parameter and automated QC that sends doubtful fits to a person.

● Detection of plate, lot and batch effects across a dataset before it is delivered, and control charting of reference standards so drift shows up over weeks as well as within a run.

● The data visualisation that accompanies each delivered result, showing it with its uncertainty, QC and batch structure, so anyone can understand and check it without asking us.

Who you are

You have worked with messy experimental data and built things other people relied on. That might have been in a biotech or pharma team, an academic lab, a core facility, or a data-heavy field outside biology where you have since picked up the biology. Nobody arrives with every part of this role. If you are strong in most of it and quick to learn the rest, we want to hear from you.

You are happy to get your hands dirty on the data. Much of the job is careful, repeated work on real instrument output, and we want someone who takes pride in doing that well. You write code other people can run, and you care how a result reads to the person receiving it.

Seniority is not what we are screening for. We are open to someone early in their career and to someone with many years behind them. What we look for is depth and judgement on experimental data.

We do not hire people into boxes. This role will stretch past its description, sometimes into the lab and sometimes into conversations beyond the team. Early hires here are expected to pick up what is in front of them.

Must have

● Building and running production data and ML pipelines, ideally on scientific or lab data, including data modelling, validation and reproducible, versioned workflows.

● Handling the specifics of biological data: plate layouts, sample and reagent lot lineage, and protein and DNA construct sequences with the standard tools for working with them.

● Biostatistics for experimental data, including experimental design, replicates and controls, curve fitting, and separating batch effects from biology.

● Statistical process control and anomaly detection, so assay drift and instrument variability are caught early.

● Data visualisation for people outside your team, presenting results with their uncertainty and QC so they can understand and check them themselves.

● Fluency in Python and its scientific stack, including pandas, NumPy and SciPy, and in working with coding agents.

● SQL and relational schema design for samples, plates and results.

● Turning scientists’ questions into requirements, and writing down the assumptions behind them.

Nice to have

● Hands-on analysis of instrument output from a lab, especially trace-based data such as SPR or HPLC.

● R for statistical modelling, such as dose-response fitting or mixed models for plate and batch effects.

● Containers and cloud infrastructure, such as Docker, Kubernetes and AWS.

● Working with data from LIMS, electronic lab notebooks or lab data standards.

● Current protein language models and structure prediction, such as ESM C, Boltz-2, Chai, OpenFold3 or AlphaFold 3.

Why this is unusual

Most assay data analysis happens far from the lab, on files someone exported weeks earlier, with little record of how the experiment ran. Here you sit next to the lab, the conditions of every run are recorded, and your analysis is delivered alongside the data.

The work also changes under you. The lab starts with people running the instruments and becomes automated over time, and the analysis you first build by hand becomes the checks that run on every result.

Some people find that energising; some find it outside their lane. It’s worth knowing in advance which one you are.

How we work

The role is based in London, at our lab and office in King’s Cross. Time in person with the scientists matters for this work. The wider team is distributed and travels, so you will work with people who are not always in the building.

We work in the open. Decisions are written down and work happens in Slack. The whole company meets once a week at the start of the week, and there is a regular offsite. Each person owns their area and makes the call inside it, and we will expect that of you early rather than late.

UK benefits are 30 days of annual leave plus public holidays, a pension with a 10% employer contribution, and Bupa private health cover. More is added as the team grows.

The team you will join

You will report to the intelligence lead. You will also work closely with the founders and with the protein characterisation scientists whose data you will analyse.

We are a small, elite team, weighted towards science and software. You will be the first person on the intelligence team whose focus is the analysis and ML infrastructure behind our results.

Our process

Our process has four stages. A short screening call about the role and the practicalities. A behavioural interview about how you work and what you value, run from the same template for every candidate. A technical stage built on a real analysis problem from our lab, not a puzzle. You spend a few hours writing up how you would approach it, using the tools you would use on the job, coding agents included, with a short note on the decisions you made. We care about judgement more than hours, so a focused answer beats an exhaustive one. We then go through it with you in person. Then references, aimed at whatever we still want to understand.

If you are not sure whether you are a fit, apply anyway. We would rather read it and decide.

We are an equal opportunity employer. We make hiring decisions on merit, scope-fit, and the strength of the working relationship we expect to build with each hire. Applications welcome from candidates of any background.

Similar jobs

Apply for this job