Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education & Training jobs
  • Remote Healthcare & Nursing jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact mahmoud@relomote.com · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Mistral logo

Research Engineer, ML Platform

Mistral
Posted 5 hours ago
📦Relocation support
🇺🇸United States
📁Data & Analytics
Is this job info correct?

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector—co-creating customized AI systems that they can run on their terms. We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited. The Role This role focuses on building and operating the ML platform that powers large-scale training, evaluation, and batch inference at Mistral AI. You will develop the infrastructure that enables researchers and engineers to run distributed GPU workloads reliably across clusters, hardware types, and regions. You will work across the full ML lifecycle, from workload scheduling and capacity management to platform APIs, observability, and production operations. You will take ownership of critical systems and help turn complex infrastructure into reliable, self-service capabilities. What You Will Do Build the ML Platform: Develop services, APIs, controllers, and tooling for training, evaluation, fine-tuning, and batch inference. Orchestrate GPU Workloads: Build systems for queueing, admission control, quotas, priorities, preemption, and topology-aware placement. Manage Compute Capacity: Improve how heterogeneous GPU resources are provisioned, allocated, and utilized across clusters. Enable Multi-Cluster Execution: Place workloads based on capacity, data locality, hardware requirements, and organizational priorities. Improve Researcher Experience: Create self-service workflows that make distributed workloads easy to launch, observe, debug, and reproduce. Optimize Performance: Improve GPU utilization, scheduling latency, workload startup time, throughput, and infrastructure efficiency. Build for Reliability: Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads. Operate What You Build: Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure. What We're Looking For Have 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field. Are proficient in Python or Go and comfortable working with production-grade distributed systems. Have strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management. Understand technologies such as Kueue, Karpenter, Volcano, and Kyverno, and the problems they address in workload scheduling, provisioning, and policy enforcement. Understand distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference. Are familiar with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking. Understand concepts such as quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission. Can diagnose performance and reliability problems across software, orchestration, networking, storage, and hardware. Care about developer experience and enjoy turning complex infrastructure into simple, reliable interfaces. Thrive in an ambiguous, fast-moving environment shaped by frontier AI research. What We Offer We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks. For the most up-to-date details on benefits available in your location, please refer to our Benefits page . Privacy Policy Your privacy matters to us. You can learn more about how we handle your personal data in our Applicant Privacy Policy .

Similar jobs

Similar jobs

Pluralis Research logo

Machine Learning Engineer - ML Training Platform

Pluralis Research

🌍Australia, United States3 days ago
Agility logo

Senior Software Engineer, AI/ML Platform

Agility

🇺🇸United States2 weeks ago
Hadrian Automation logo

ML Platform Engineer

Hadrian Automation

🇺🇸United States4 weeks ago
Sola logo

Software Engineer, ML Platform

Sola

🇺🇸United StatesMay 28, 2026, 3:07 AM UTC
RI

Artificial Intelligence Engineer

RPL International

🇺🇸United States3 hours ago
ME

Associate Director, AI/ML Engineering

Merck

🇺🇸United States3 hours ago