Member of Technical Staff, ML Systems
- Salary
- $180K–$230KUSD per year
- Moves you to
- United States
- Support
- Visa sponsorship
- Posted
- Oct 1, 2026
About the Role
This is a systems-focused ML infrastructure role at an early-stage AI startup rebuilding the training and inference stack for world models, from GPU kernels all the way through distributed serving. You will own speed and efficiency across the full stack, working closely with a small founding team that covers distributed systems, kernel optimization, cloud infrastructure, and research. If you would rather make a video model ten times faster than train one, this role is for you.
What You'll Do
Optimize GPU and system performance for training and inference across image, video, and world-model workloads.
Profile and remove bottlenecks at the kernel, memory, system, and cluster level using Nsight and related tooling.
Write low-level CUDA and Triton optimizations on production code paths.
Build distributed inference and training engines for diffusion models across multiple GPUs and nodes.
Own communication performance covering NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving.
Maintain benchmarking and regression harnesses to ensure performance gains hold in production.
What We're Looking For
1 or more years of full-time, hands-on work on inference or training performance: GPU kernels, runtime, or distributed execution.
Authorship of core features in an inference or training framework such as vLLM, SGLang, TensorRT-LLM, Megatron, or similar (not deployment or integration work).
Hands-on experience writing CUDA, CUTLASS, Triton, or PTX kernels on NVIDIA GPUs.
Strong CS fundamentals from a degree in computer science or a related quantitative field.
Experience optimizing diffusion, video, image, or other multimodal workloads is a strong plus.
Background in compiler engineering with MLIR, LLVM, or codegen is a plus.
Familiarity with multi-GPU or multi-node communication technologies such as NCCL, RDMA, InfiniBand, or RoCE is a plus.
Experience with AMD, FPGA, or custom-accelerator platforms is a plus.
Compensation & Benefits
Base salary: $180,000 to $230,000 USD annually, plus equity. Visa sponsorship is available.
Location
On-site in Menlo Park, California, United States.