Member of Technical Staff, ML Systems
- Salary
- $180K–$230KUSD per year
- Moves you to
- United States
- Support
- Visa sponsorship
- Posted
- Sep 30, 2026
About the Role
This role sits at the core of an early-stage AI infrastructure team rebuilding the training and inference stack for world models, from GPU kernels all the way through distributed serving. You will own speed and efficiency across that entire stack, working directly alongside founding engineers who cover distributed systems, kernel optimization, cloud infrastructure, and research. If you would rather make a video model ten times faster than train one, this is the right place.
What You'll Do
Optimize GPU and system performance for image, video, and world-model training and inference workloads.
Profile and remove bottlenecks at the kernel, memory, system, and cluster level using Nsight and related tooling.
Write low-level CUDA and Triton optimizations on production code paths.
Build distributed inference and training engines for diffusion models across multiple GPUs and nodes.
Own communication performance across NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving topologies.
Maintain benchmarking and regression harnesses so performance gains hold in production over time.
What We're Looking For
1 or more years of full-time, hands-on work on inference or training performance: GPU kernels, runtime, or distributed execution.
Demonstrated authorship of core features in an inference or training framework such as vLLM, SGLang, TensorRT-LLM, Megatron, or a comparable project, not deployment or integration work.
Experience writing CUDA, CUTLASS, Triton, or PTX kernels on NVIDIA GPUs.
Strong computer science fundamentals from a degree or equivalent rigorous background in a quantitative field.
Familiarity with multi-GPU or multi-node communication technologies such as NCCL, RDMA, InfiniBand, or RoCE.
Experience optimizing diffusion, video, image, or other multimodal workloads is a strong plus.
Background as a compiler engineer working with MLIR, LLVM, or codegen is a plus.
Experience with AMD, FPGA, or custom-accelerator platforms is a plus.
Compensation & Benefits
Salary: $180,000 to $230,000 per year
Visa sponsorship is available.
Location
On-site in Menlo Park, California, United States.