SC

Site Reliability Engineer – GPU/HPC Infrastructure (Remote – MENA)

Saturn Cloud
Posted 1 hour ago
Middle EastRemoteEngineering & Development
Is this job info correct?

This is a fully remote role open to candidates across the MENA region


About Saturn Cloud

Saturn Cloud builds infrastructure for running AI, machine learning, and data workloads at scale. Our platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure.


We’re looking for a GPU/HPC-focused Site Reliability Engineer based in the MENA region to help operate and troubleshoot the large-scale GPU infrastructure supporting Saturn Cloud Token Factory.


This role is focused on the infrastructure below and around the Kubernetes layer. You’ll serve as a technical escalation point for GPU health, NVIDIA software, high-performance networking, topology, and distributed GPU performance issues affecting production inference workloads.

The ideal candidate comes from GPU infrastructure, HPC, AI infrastructure, neocloud, hyperscaler, or large-scale ML platform operations rather than traditional application support.


We value engineers who actively use modern AI coding agents to improve the speed and quality of their engineering and operational work.


What You’ll Do

  • Diagnose and resolve production issues affecting large-scale GPU inference infrastructure
  • Troubleshoot NVIDIA datacenter GPUs, drivers, CUDA compatibility, and GPU container runtimes
  • Investigate GPU health issues, including Xid errors and hardware or driver failure modes
  • Diagnose PCIe, NUMA, GPU placement, NVLink, and NVSwitch issues
  • Troubleshoot multi-GPU and multi-node workloads
  • Investigate high-performance networking and distributed communication failures
  • Distinguish application and inference-runtime issues from GPU, fabric, topology, node, driver, or hardware failures
  • Work with Kubernetes-based GPU workloads and NVIDIA GPU Operator/device plugins
  • Use production observability and GPU metrics to diagnose reliability and performance issues
  • Work directly with GPU-cloud and infrastructure providers when incidents require hardware or fabric investigation
  • Produce clear technical evidence showing where the failure occurs and what the appropriate infrastructure team needs to investigate


What We’re Looking For


Core Skills

  • Hands-on experience using AI coding agents and agentic development tools to accelerate engineering, debugging, automation, and operational workflows
  • Deep Linux systems debugging experience
  • NVIDIA datacenter GPU administration
  • NVIDIA driver installation, upgrades, and troubleshooting
  • Strong understanding of CUDA and driver compatibility
  • NVIDIA Container Toolkit/runtime
  • Experience with NVML and nvidia-smi
  • GPU health diagnostics, including Xid errors and common hardware/driver failure modes
  • Understanding of PCIe topology, NUMA, and GPU placement
  • Understanding of NVLink and NVSwitch fundamentals
  • Familiarity with Kubernetes GPU Operator and device plugins
  • Experience running containerized GPU workloads


High-Performance Networking


You should have strong experience with several of the following:

  • InfiniBand
  • RDMA
  • RoCE
  • NCCL
  • GPUDirect RDMA
  • NIC/GPU topology
  • NCCL testing and distributed workload diagnostics
  • Bandwidth and latency troubleshooting
  • Multi-node GPU communication failures


You should be able to determine whether a production issue originates in the application, GPU, network fabric, topology, node, or another underlying infrastructure layer.


Inference Infrastructure

You don’t need to be an ML researcher, but you should understand how modern inference workloads exercise GPU infrastructure.


Relevant experience includes:

  • vLLM, NVIDIA Dynamo, Triton, or comparable inference runtimes
  • Tensor and pipeline parallelism
  • Model loading and GPU memory consumption
  • KV cache
  • Continuous batching
  • GPU and memory utilization
  • Out-of-memory diagnosis
  • Multi-GPU and multi-node inference
  • Basic inference performance analysis, including throughput, latency, and GPU saturation


Additional Skills

  • Kubernetes troubleshooting sufficient to independently investigate GPU workloads inside a cluster
  • Prometheus/Grafana and NVIDIA DCGM metrics
  • Bash and Python
  • containerd/Docker
  • Experience with bare-metal GPU clusters or GPU cloud infrastructure
  • Familiarity with B200/B300/H200/H100-class systems is highly desirable


Strong Pluses

  • NVIDIA DCGM
  • NVIDIA Dynamo
  • KAI Scheduler or Grove
  • Spectrum-X
  • Mellanox/ConnectX networking
  • OFED/DOCA
  • Slurm or HPC cluster administration
  • Kubernetes-based GPU clouds
  • Large GPU fleet operations
  • GPU burn-in, qualification, and health-check tooling


What Success Looks Like


When an inference workload becomes unavailable or significantly slower, you can determine whether the problem is caused by the inference runtime, GPU memory pressure, a failed GPU, a driver/kernel interaction, PCIe/NVLink/NVSwitch topology, NCCL, RDMA/InfiniBand fabric, or the underlying node.


You can collect enough evidence to distinguish a Saturn Cloud software issue from an operator infrastructure or hardware problem and communicate precisely what an infrastructure provider needs to investigate.


You’re the person the team turns to when Kubernetes looks healthy, but the GPU workload isn’t.


Why Saturn Cloud

  • Remote-first culture with a high-trust, high-ownership environment
  • Work directly with cutting-edge GPU and AI infrastructure
  • Solve complex production problems spanning hardware, networking, Kubernetes, and modern inference systems
  • Help shape the reliability and operational practices behind large-scale AI workloads


Compensation & Benefits

  • Competitive salary
  • Flexible PTO
  • Fully remote position

Similar jobs