Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
SI

Cluster Engineer

STN Inc
Posted 5 hours ago
🌍Probably Worldwide🏠Remote📁Engineering & Development
Is this job info correct?

Position Summary We are seeking a highly experienced AI Infrastructure Engineer to architect, deploy, optimize, and operate large-scale GPU clusters supporting state-of-the-art AI training and inference workloads. This is a deeply technical role focused on maximizing cluster efficiency, scalability, and performance across the entire AI stack—from GPU hardware and high-speed networking to distributed training frameworks and inference optimization. The ideal candidate has built GPU clusters from the ground up, tuned distributed training environments, optimized large-scale inference deployments, and understands how every layer of the infrastructure contributes to application performance. Responsibilities Design, deploy, and optimize multi-node GPU clusters for AI training and inference workloads. Tune distributed training environments to maximize GPU utilization, throughput, and scaling efficiency. Optimize inference clusters for maximum token generation throughput, low latency, and high GPU utilization. Build and support production AI infrastructure running hundreds to thousands of GPUs. Analyze and eliminate performance bottlenecks across compute, networking, storage, and software layers. Perform NCCL benchmarking, analysis, and tuning to achieve optimal collective communication performance. Design and optimize GPU networking using InfiniBand or RoCE v2, including RDMA, congestion management, topology awareness, and QoS. Configure and tune distributed AI software stacks including: PyTorch NCCL CUDA UCX MPI Slurm Pyxis/Enroot Optimize GPU scheduling and resource allocation for both training and inference environments. Develop repeatable benchmarking and validation processes for new hardware, firmware, drivers, and software releases. Identify performance regressions and troubleshoot distributed training issues at scale. Optimize storage architectures for AI workloads, including checkpointing, dataset streaming, and high-performance parallel I/O. Work closely with ML engineers to improve training scalability and inference efficiency. Create automation to deploy, validate, benchmark, and monitor GPU clusters. Evaluate emerging AI infrastructure technologies and recommend improvements to platform architecture. Required Qualifications 7+ years designing or operating large-scale Linux infrastructure. 5+ years supporting production GPU clusters for AI or HPC workloads. Demonstrated experience building multi-node GPU training environments from the ground up. Deep expertise with distributed PyTorch training. Extensive experience troubleshooting and optimizing NCCL communications. Strong understanding of distributed AI communication patterns, including: AllReduce ReduceScatter AllGather Broadcast Point-to-point communications Experience benchmarking distributed training using tools such as: nccl-tests NVIDIA DCGM Nsight Systems MLPerf (preferred) Strong understanding of GPU memory management, including: KV Cache Activation checkpointing Tensor Parallelism Pipeline Parallelism Data Parallelism Experience optimizing LLM inference throughput, including: Tokens/sec optimization Batch sizing Continuous batching KV cache tuning Memory bandwidth optimization Experience tuning CUDA, NCCL, UCX, and MPI for maximum distributed performance. Expert-level Linux systems administration skills. Experience with Slurm workload manager. Experience using Pyxis and Enroot for containerized GPU workloads. Strong scripting skills using Python and Bash. Technical Expertise AI Frameworks PyTorch CUDA NCCL Triton (preferred) TensorRT-LLM (preferred) Cluster Scheduling Slurm Pyxis Enroot GPU Networking Strong understanding of: InfiniBand RoCE v2 RDMA GPUDirect RDMA GPUDirect Storage UCX MPI Network topology optimization Congestion control QoS ECN/PFC High-speed Ethernet (200/400/800 GbE) Storage Experience designing or tuning storage for AI workloads, including: Parallel file systems Distributed storage Object storage NVMe Checkpoint optimization Dataset staging GPUDirect Storage Storage bandwidth optimization Metadata performance Performance Engineering Experience with: NCCL benchmarking Multi-node scaling analysis GPU utilization optimization Communication/computation overlap NUMA optimization CPU affinity PCIe topology GPU topology (NVLink/NVSwitch) Memory bandwidth analysis End-to-end performance profiling Preferred Qualifications Experience deploying AI workloads on Kubernetes. Experience with NVIDIA GPU Operator. Experience with Kubernetes batch scheduling (Volcano, Kueue, Run:ai, etc.). Experience with distributed inference platforms such as vLLM, TensorRT-LLM, or SGLang. Experience with NVIDIA DGX SuperPOD or similar large-scale GPU deployments. Familiarity with MLPerf benchmarking. Experience deploying monitoring solutions such as Prometheus, Grafana, and DCGM Exporter. Experience automating infrastructure using Ansible, Terraform, or similar tools. Experience working in cloud GPU environments (AWS, Azure, GCP) in addition to bare metal.

Similar jobs

Similar jobs

Sembi logo

IT Platform Engineer (Data Lake and AI Gateway)

Sembi

🌍Probably Worldwide4 hours ago
AO Globe Life logo

Remote Leadership Development

AO Globe Life

🌍Probably Worldwide4 hours ago
Elite Clinical Network logo

IT Support Specialist

Elite Clinical Network

🌍Probably Worldwide4 hours ago
Jtl Software Gmbh logo

Director Customer Service & Support (m/w/d)

Jtl Software Gmbh

🌍Probably Worldwide4 hours ago
株式

グローバル経理 ※経理未経験OK

株式会社ウェザーニューズ

🌍Probably Worldwide4 hours ago
DistantJob logo

Bilingual (French/English) Back-Office Coordinator

DistantJob

🌍Probably Worldwide4 hours ago