Member of Technical Staff, ML Systems
- Salary
- $180K–$230K
- Moves you to
- United States
- Support
- Visa sponsorship
- Posted
63,684 relocation jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
About the Role
Own performance across the training and inference stack for image, video, and world-model workloads, from GPU kernels to distributed systems. You will work with an engineering and research team to make these workloads faster and more efficient.
What You'll Do
Optimize training and inference performance across GPU kernels, memory, systems, and clusters.
Profile workloads with Nsight and related tools, identify bottlenecks, and implement low-level optimizations in CUDA or Triton.
Design and improve distributed training and inference engines for diffusion models.
Improve communication across GPUs and nodes using technologies such as NCCL, RDMA, InfiniBand, or RoCE.
Build benchmarks and regression tests to ensure performance improvements hold in production.
Collaborate on hardware-aware kernel, runtime, and model design.
What We're Looking For
At least 2 years of experience in deep learning training or inference systems, or distributed systems, including at least 1 year of hands-on ML systems or GPU performance work.
Experience authoring core features in an inference or training framework such as vLLM, SGLang, TensorRT-LLM, or Megatron, rather than focusing only on deployment or integration.
Direct GPU kernel optimization experience on NVIDIA GPUs using CUDA, Triton, CUTLASS, or PTX.
Experience with PyTorch, Nsight, and multi-GPU or multi-node communication.
Strong machine learning and computer science fundamentals, and a degree in computer science or a related quantitative field.
Experience optimizing diffusion, video, image, or other multimodal workloads is valuable.
Compensation & Benefits
Annual salary range: $180,000 to $230,000 USD. Visa sponsorship is available.
Location
On-site in Menlo Park, California, United States.