Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Cerence logo

Senior Principal AI Engineer

Cerence
Posted 5 hours ago
🇺🇸United States🏠Remote📁Engineering & Development
Is this job info correct?

A Moving Experience. What You Will Work On Design and operate distributed training systems for large neural networks (autoregressive, diffusion , State Space Models etc. ) across GPU clusters Optimise multi ‑ node, multi ‑ GPU execution to maximize throughput and utilization Diagnose & resolve bottlenecks across compute, memory, and netwo rk Improve training stability and fault tolerance at scale Partner with research and applied ML teams to productioni z e large ‑ model training pipelines Core Responsibilities Distributed Training Infrastructure Build and optimize GPU cluster orchestration using: Slurm Kubernetes Ray RunAI Ensure efficient scheduling, isolation, and fairness across training workloads Communication & Networking Optimize and debug distributed communication using: NCCL RDMA InfiniBand NVLink Minimize networking bottlenecks that dominate end ‑ to ‑ end training time Training Frameworks Scale large-model training using: PyTorch Distributed Megatron ‑ LM DeepSpeed Own multi ‑ node launch configurations, failure recovery, and performance tuning Memory & Performance Optimization Apply advanced memory optimization techniques: Activation checkpointing ZeRO (Stage 1–3) and offload strategies Balance compute, memory, and communication to push model size and batch scale What Success Looks Like GPU utilization consistently stays high (>80–90%) Training scales cleanly from single node to dozens or hundreds of GPUs Communication overhead is minimized and predictable Large training jobs run stably for days or weeks without failure New models can be trained faster, larger, and more reliably than before Required Experience & Skills Strongly Required Deep hands ‑ on experience with distributed systems or ML systems Experience running large ‑ scale workloads on GPU clusters Production experience with PyTorch distributed training Strong understanding of parallelism strategies (data, tensor, pipeline parallelism) Low ‑ level understanding of GPU communication and networking Critical Technical Skills GPU orchestration: Slurm , Kubernetes, Ray, RunAI Communication libraries: NCCL, RDMA, InfiniBand, NVLink Training frameworks: PyTorch Distributed, Megatron ‑ LM , DeepSpeed Memory optimi s ation : activation checkpointing, ZeRO offload techniques Common Problems You’ll Be Solving Many teams fail at scale because: GPU utilization is low despite large clusters Networking and communication dominate training time Training jobs crash or become unstable at large scale You will be explicitly focused on eliminating these failure modes. Ideal Background This role is a strong fit for individuals who have worked as: ML Systems Engineer Distributed Systems Engineer AI Infrastructure Engineer HPC Engineer transitioning into ML Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience. Why This Role Matters Without robust distributed training infrastructure, progress on large models stalls. This role directly enables: Larger models Faster iteration cycles More reliable research-to-production pipelines You will be building the foundation that makes large ‑ scale AI possible. Cerence Inc. (Nasdaq: CRNC and www.cerence.com ) is the global industry leader in creating unique, moving experiences for the automotive world. Spun out from Nuance in October 2019, Cerence is a new, independent company that has quickly gained traction as a leader in the automotive voice assistant space, working with all of the world’s leading automakers – from Ford and Fiat Chrysler to Daimler, Audi and BMW to Geely and SAIC – to transform how a car feels, responds and learns. Its track record is built on more than 20 years of industry experience and leadership and more than 500 million cars on the road today across more than 70 languages. As Cerence looks to the future and continues an ambitious growth agenda, we need someone to join the team and help build the future of voice and AI in cars. This is an exciting opportunity to join Cerence’s passionate, dedicated, global team and be a part of meaningful innovation in a rapidly growing industry. EQUAL OPPORTUNITY EMPLOYER Cerence is firmly committed to Equal Employment Opportunity (EEO) and to compliance with all federal, state and local laws that prohibit employment discrimination on the basis of age, race, color, gender, gender identity, gender expression, sex, sex stereotyping, pregnancy, national origin, ancestry, religion, physical or mental disability, medical condition, marital status, citizenship status, sexual orientation, protected military or veteran status, genetic information and other protected classifications. Cerence Equal Employment Opportunity Policy Statement. All prospective and current Employees need to remain vigilant when it comes to executing security policies in the workplace. This includes: - Following workplace security protocols and training programs to familiarize with the ways to maintain a safe workplace. - Following security procedures to report any suspicious activity. - Having respect for corporate security procedures to allow those procedures to be effective. - Adhering to company's compliance and regulations. - Encouraging to follow a zero tolerance for workplace violence. - Basic knowledge of information security and data privacy requirements (e.g., how to protect data & how to be handling this data). - Demonstrative knowledge of information security through internal training programs.

Similar jobs

Similar jobs

Leidos logo

Senior AWS Generative AI Solution Engineer

Leidos

🇺🇸United States3 hours ago
Initiate Government Solutions logo

AI/ML Engineer

Initiate Government Solutions

🇺🇸United States4 hours ago
Globo Language Solutions logo

AI/ML Engineer

Globo Language Solutions

🇺🇸United States4 hours ago
Choicehotels logo

Software Engineer 2 - AI Operations & Enablement

Choicehotels

🇺🇸United States4 hours ago
Nvidia logo

Senior AI Infrastructure Software Engineer - DGX Cloud

Nvidia

🇺🇸United States4 hours ago
Nvidia logo

Senior AI Developer Technology Engineer, Financial Sector

Nvidia

🇺🇸United States4 hours ago