Member of Technical Staff, ML Systems
- Salary
- $180K–$230KUSD per year
- Moves you to
- United States
- Support
- Visa sponsorship
- Posted
- Oct 2, 2026
About the Role
This role sits at the heart of an early-stage AI/ML infrastructure team rebuilding the training and inference stack for world models, from GPU kernels all the way through distributed serving. You will own speed and efficiency across the full stack, working closely with a small founding team covering distributed systems, kernel optimization, cloud infrastructure, and research. If you would rather make a video model ten times faster than train one, this is the place for you.
What You'll Do
Optimize GPU and system performance for image, video, and world-model training and inference workloads.
Profile and remove bottlenecks at the kernel, memory, system, and cluster level using Nsight and related tooling.
Write low-level optimizations in CUDA and Triton on production code paths.
Build distributed inference and training engines for diffusion models across multiple GPUs and nodes.
Own communication performance across NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
Maintain performance gains through benchmarking and regression harnesses so improvements hold in production.
What We're Looking For
2 or more years of experience in deep-learning inference, training systems, or distributed systems.
Hands-on experience writing low-level GPU optimizations with CUDA or Triton.
Background optimizing diffusion, video, image, or other multimodal workloads.
Experience with multi-GPU or multi-node communication technologies such as NCCL, RDMA, InfiniBand, or RoCE.
Experience as a compiler engineer working with MLIR, LLVM, or codegen is a strong plus.
Familiarity with AMD, FPGA, or custom-accelerator platforms is a strong plus.
Comfort working with PyTorch at a systems level.
Compensation & Benefits
Base salary: $180,000 to $230,000 USD annually. Visa sponsorship is available.
Location
On-site in Menlo Park, California, United States.