Our client is a hyper-growth, heavily backed startup that is revolutionizing the frontier of healthcare AI to radically improve drug discovery and solve cancer. Having recently secured a $100M raise, they are building the foundational data layer and post-training simulation environments where the next generation of medical AI models learn. They are seeking a powerhouse Research Engineer to own the full reinforcement learning loop and build the crucial evaluation and reward layer for frontier AI models using real-world clinical and genomic context.
Role & Impact
- Build comprehensive RL environments, encompassing task design, action spaces, tool interfaces, verifiers, and evaluation harnesses.
- Run sophisticated post-training experiments on language models and agents utilizing techniques like SFT, RLVR, RLHF/RLAIF, and reward modeling.
- Train and evaluate multi-step agents to navigate complex patient histories and unstructured medical data to capture correctness in clinical reasoning.
- Analyze model failures to continuously improve the data, rewards, and feedback signals that dictate what these systems learn.
Essential Skills
- 1+ years of highly technical experience focusing on reinforcement learning, agent environments, LM post-training, or related ML systems.
- Exceptional judgment around data quality, capable of assessing signal fidelity, coverage, and clinical relevance to support meaningful training tasks.
- Strong ability to build scalable pipelines, debug complex training runs within large ML codebases, and move rapidly from research concepts to working prototypes.
Location: San Francisco, CA (In-person, with relocation and joining bonuses available)
Compensation: Up to $350,000 base + up to 1.0% equity + comprehensive benefits including unlimited AI tool budgets
If this aligns with your background, reply and Goliath Partners will be in touch!