About Modular At Modular, we’re on a mission to revolutionize AI infrastructure by systematically rebuilding the AI software stack from the ground up. Our team, made up of industry leaders and experts, is building cutting-edge, modular infrastructure that simplifies AI development and deployment. By rethinking the complexities of AI systems, we’re empowering everyone to unlock AI’s full potential and tackle some of the world’s most pressing challenges. If you’re passionate about shaping the future of AI and creating tools that make a real difference in people’s lives, we want you on our team. You can read about our culture and careers to understand how we work and what we value. About the role: Modular is building a next-generation AI infrastructure platform that unifies the many application frameworks and hardware backends, simplifying deployment for AI production teams and accelerating innovation for AI researchers and hardware developers. We are looking for a Developer Advocate to evangelize the MAX Platform's inference and serving capabilities with our user base and developer community. This involves creating technical content such as user guides and blog posts as well as giving talks at conferences, leading workshops, all with the goal of enabling our community of builders deploying models in production. Join our world-leading product team and be part of redefining how AI infrastructure is built and deployed. LOCATION: Candidates based in the US or Canada are welcome to apply. To support growth and collaboration, those in earlier career stages work in a hybrid capacity at our Los Altos, CA. More senior staff can work out of our office in Los Altos, CA or remotely from home. Onboarding for new hires is conducted in-person in our Los Altos, CA office. Additionally, this role requires travel to conferences and developer events, which may be as often as once per month, as well as travel for team and company events (typically 2-4 times per year). What you will do: Provide support and respond to questions from customers evaluating MAX for inference and serving, including some of the world's largest corporations. Establish, run and publish benchmarks comparing MAX to serving frameworks like vLLM, Triton Inference Server, and TensorRT-LLM. Foster an inclusive and welcoming environment for ML engineers and practitioners deploying models with MAX. Collaborate with engineering and product teams to create tutorials, video guides, and examples demonstrating how the MAX Platform handles inference and serving workloads efficiently on both CPUs and GPUs. Write blog posts and other educational content that inform prospective and current customers about MAX's inference and serving performance and functionality. Engage with the community across GitHub, Discord, Twitter/X, and LinkedIn by facilitating discussions, answering questions, and providing support. Act as a voice for the developer community internally, feeding inference and serving feedback back to engineering and product teams, shaping the future of the MAX Platform. Represent Modular at conferences, summits, and industry events through presentations, panel discussions, and networking with industry professionals. Contribute to Modular's broader developer relations, product, and marketing strategies. What success looks like after 6 months You've published a steady cadence of inference and serving content that developers reference when deploying models with MAX. Your benchmarks and comparisons against tools like vLLM, Triton Inference Server, and TensorRT-LLM get cited in community discussions. Code examples and cookbooks you've built get linked in forum and Discord answers by other community members. You've identified and closed content gaps that were blocking adoption of MAX for production inference. What you bring to the table: Demonstrable experience creating technical content for developer audiences. Send us a portfolio: blog posts, tutorials, videos, docs, or courses. You understand the ML inference stack, including model serving architectures, GPU acceleration, and how MAX compares to vLLM, Triton Inference Server, and TensorRT-LLM. Strong Python skills; systems programming experience (C++, Rust, or similar) is an advantage. You learn new tools fast and produce accurate content quickly. A feature ships Tuesday, your tutorial goes out Thursday. You can record, edit, and publish a technical video without a production team. Clear audio, good pacing, technically correct, not necessarily polished. You write well. You explain complex ideas without losing precision, and you cut the filler. You plan your own content calendar because you're in the community and know where developers get stuck. A growth and leadership mindset, with a collaborative attitude that seeks to learn more from our customers, team members, and the broader market. Helpful, but not required Experience programming GPUs using CUDA or ROCm. Familiarity with the Mojo 🔥 programming language and MAX AI framework. Familiarity with open source software development practices and communities. Experience producing high-quality videos covering technical topics. What Modular brings to the table: Amazing Team. We are a progressive and agile team with some of the industry’s best engineering and product leaders. World-class Benefits. In order to attract the best, we need to offer the best. Premier insurance plans, up to 5% 401k matching, flexible paid time off, and more are available to you! Please note that specific benefit packages may vary based on your location. Competitive Compensation. We offer very strong compensation packages, including stock options. We want people to be focused on their best work and believe in tailoring compensation plans to meet the needs of our workforce. Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as different cities. Traveling 2-4 times a year is expected for all roles. Working at Modular will enable you to grow quickly as you work alongside incredibly motivated and talented people who have high standards, possess a growth mindset, and a purpose to truly change the world. The estimated base salary range for this role to be performed in the US is $150,200.00 - $225,400.00 USD . The estimated base salary range for this role to be performed in Canada is $111,500.00 - $167,300.00 CAD . The salary for the successful applicant will depend on a variety of permissible, non-discriminatory job-related factors, which include but are not limited to education, training, work experience, business needs, or market demands. This range may be modified in the future. The total compensation for a candidate will also include annual target bonus, equity, and benefits, with equity making up a significant portion of your total compensation. For candidates who fall outside of the listed requirements, we nevertheless encourage you to apply as we may have openings that are lower/higher level than the ones advertised.
Developer Advocate, MAX Inference & Serving (Copy)
Modular
Inference Infrastructure Engineer, Serving
Elorian
Senior Engineer II, Inference Engine - Serving Engine
DigitalOcean
Product Manager - AI Inference & Model Serving
Mirantis
Machine Learning Engineer, Inference & Serving (Speech LLM) - San Francisco
Plaud
Director of Software Engineering
Aerostrat