Engineering Manager, ML Performance Optimization
- Salary
- $276K–$343K
- Hiring from
- United States
- Work type
- Hybrid
- Posted
Show job descriptionHide job description
In this role, you will:
-
Vision: Develop and execute a strategic vision and roadmap for ML Training and Inference Performance Optimization, ensuring scalability, reliability, and performance to support autonomous driving.
-
Technical acumen: Lead the design, implementation, and operation of a robust and efficient ML platform to enable the training, validation, serving, optimization and monitoring of ML models.
-
ML Performance Optimization: Drive end-to-end performance optimization for large-scale model training and inference, including distributed training efficiency, GPU utilization, memory and communication optimization, model compression (quantization, pruning, distillation), and low-latency on-vehicle inference that meets strict real-time and compute budgets.
-
Hiring: Attract, hire, and inspire a diverse world-class engineering team, fostering a culture of innovation, collaboration, and excellence.
-
Partnership: Collaborate closely with cross-functional teams, including ML researchers, software engineers, data engineers, and hardware engineers, to define requirements and align on architectural decisions.
-
Mentorship: Enable engineers on the team to grow their careers by providing the right opportunities and clear, timely feedback.
Qualifications
- 8+ years of relevant experience, including 3+ years of management experience managing engineers.
- Strong technical background in ML performance optimization, such as distributed training strategies (data, tensor, pipeline parallelism, FSDP/ZeRO), mixed-precision training, kernel-level optimization (CUDA, Triton), compiler stacks (torch.compile, XLA, TVM), quantization, and profiling/benchmarking across GPU and embedded accelerators.
- Experience building user-friendly ML Infrastructure that enabled large-scale model training and high-throughput, low-latency serving use cases.
- Experience with training frameworks like PyTorch, JAX, etc., leveraging GPUs for distributed model training.
- Experience with GPU-accelerated inference using TensorRT, Ray Serve, or similar frameworks.
- Proven track record of extensive cross-functional collaboration, partnering with research, product, hardware, and platform teams to align priorities, influence technical direction, and deliver measurable performance improvements across organizational boundaries.