Compute Platform Lead | Frontier AI Research Lab - New York, SF or London
- Moves you to
- United States
- Support
- Visa sponsorship
- Posted
63,681 relocation jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
An extremely well-funded and innovative AI research lab training frontier-scale foundation models is seeking a Compute Platform Lead. This person will own the GPU compute layer that its largest training runs depend on.
The team runs a Kubernetes-based platform across several GPU cloud providers. That makes for hard systems problems: scheduling across multiple clouds, node health, and debugging performance at scale. You will lead the team that builds and runs this layer, and you will stay hands-on enough to ship code yourself.
What you will do
- Build, mentor and grow a team of strong systems engineers, scaling towards about 10.
- Own fleet reliability and availability: scheduling across clouds, cluster management, and getting ready for next-generation GPUs and much larger clusters.
- Make the architecture calls on automatic remediation, topology-aware scheduling (placing jobs by network layout), capacity planning, hardware debugging, and fleet-wide monitoring and benchmarking.
- Work with the training teams to design fault tolerance, node health checks and remediation together.
- Own vendor relationships, including negotiating and running major compute deals.
- Over time, take on multi-cloud storage, petabyte-scale data replication and GPU-to-GPU network performance.
What they are looking for
- A track record of building and leading infrastructure or systems teams while staying technical.
- Deep systems engineering experience with how whole clusters behave and how to maintain them.
- Strong coding ability and the credibility to earn a senior team's technical trust.
- Depth in at least one of orchestration, storage or GPU hardware. GPU knowledge beyond standard Kubernetes (e.g. NCCL, NVIDIA's GPU communication library) is a plus.
- Experience with high-performance storage across data centres and with checkpointing at scale is a plus.
- Commercial judgement managing vendors and closing significant deals.
What is on offer
- Top-of-market salary plus meaningful equity.
- Comprehensive health cover, generous paid parental leave and flexible time off.
- Visa sponsorship for exceptional candidates.
- A small, talent-dense team where this hire shapes how the lab scales its compute.
If you have run GPU fleets at scale and want to lead the platform under frontier training runs, message me directly or comment below. All conversations are confidential.
#Hiring #AIInfrastructure #GPU #Kubernetes #MLInfrastructure #EngineeringLeadership