Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Institute of Foundation Models logo

Senior Distributed Systems Engineer

Institute of Foundation Models
Posted May 28, 2026, 12:07 AM UTC
🛂Visa sponsorship
🇺🇸United States
📁Engineering & Development
Is this job info correct?

About the Institute of Foundation Models The Institute of Foundation Models (IFM) designs and operates ultra-scale GPU supercomputing systems to train next-generation foundation models. We believe performance, fault tolerance, and scalability are co-designed across model architecture, communication systems, runtime, and hardware topology. This role sits at the core of that effort — driving communication performance, distributed reliability, and cross-layer optimization for large-scale training workloads. The Mission We are looking for a deeply technical engineer to co-design and optimize the communication stack for large-scale distributed training, including hybrid parallelism and Mixture-of-Experts (MoE) workloads. This is not a network operations role. This is a systems-level engineering position focused on performance engineering, distributed debugging, and communication-runtime co-design. · Design and optimize expert-parallel and hybrid-parallel communication patterns · Drive high-performance hierarchical collectives for MoE workloads · Co-design runtime orchestration with communication topology awareness · Reduce tail latency and improve determinism across thousands of GPUs · Architect fault-tolerant distributed execution under real-world cluster failures Core Technical Scope · Communication-compute overlap and topology-aware collective optimization · Deep debugging of NCCL, RDMA, and custom communication layers · Hybrid expert parallel strategies in modern large-scale MoE systems · Elastic and resilient distributed job orchestration concepts · Congestion analysis and routing optimization across InfiniBand/RoCE fabrics · Microbenchmarking and performance modeling for communication-heavy workloads Expected Technical Depth · Hybrid expert parallel communication for Mixture-of-Experts training · Scaling behavior under network pressure · Distributed orchestration for elastic, large-scale training · Fault detection and recovery in distributed GPU workloads · Cross-layer bottlenecks: GPU ↔ NIC ↔ PCIe ↔ NVSwitch ↔ Fabric ↔ Scheduler Required Background · Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) · Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA · Deep familiarity with NCCL and/or UCX internals · Strong systems programming ability (C/C++, Rust, or Go) · Strong familiarity with modern model training frameworks such as PyTorch · Ability to troubleshoot and profile training performance issues related to communication bottlenecks · Ability to translate research ideas into production-grade optimizations · Experience debugging distributed hangs, desynchronization, and performance regressions What We Mean by "Hardcore" · You can explain why an communication degrades at scale and how to fix it · You have improved real cluster throughput via communication redesign · You can trace a distributed hang across ranks and identify the root cause · You are comfortable working at the boundary between hardware and runtime Application Requirements · Include a link to your GitHub (required) · Provide links to relevant distributed systems, HPC, or large-scale training projects · Include a list of publications and/or public technical reports (if applicable) · Describe the hardest distributed debugging problem you solved · Include measurable performance improvements you have delivered Academic Qualifications Master’s, or Bachelor’s + 1 year of relevant experience. Visa Sponsorship This position is eligible for visa sponsorship. Benefits Include *Comprehensive medical, dental, and vision benefits *Bonus *401K Plan *Generous paid time off, sick leave and holidays *Paid Parental Leave *Employee Assistance Program *Life insurance and disability

Similar jobs

Similar jobs

Clera logo

Senior Software Engineer, Distributed Data Systems

Clera

🇺🇸United States10 hours ago
CyberCoders logo

Sr Embedded Systems Engineer (Hardware-focused)

CyberCoders

🇺🇸United StatesYesterday
Blue River Technology logo

Principal Autonomous Systems Safety Engineer

Blue River Technology

🇺🇸United States2 days ago
Santa Clara University logo

Electrical and Computer Engineering Assistant Professor - Digital Systems

Santa Clara University

🇺🇸United States2 days ago
Verkada logo

Staff Frontend Engineer - Design Systems

Verkada

🇺🇸United States2 days ago
Twelve Labs logo

Senior GTM Systems Engineer

Twelve Labs

🇺🇸United States2 days ago