Site Reliability Engineer – Kubernetes/Linux Platform (Remote – MENA)
Saturn CloudThis is a fully remote role open to candidates across the MENA region
About Saturn Cloud
Saturn Cloud builds infrastructure for running AI, machine learning, and data workloads at scale. Our platform helps teams develop, deploy, and operate compute-intensive workloads across modern cloud and GPU infrastructure.
We’re looking for a Site Reliability Engineer based in the MENA region to help operate and troubleshoot the production infrastructure supporting Saturn Cloud Token Factory.
This is an infrastructure/SRE role rather than a traditional customer support position. You’ll be expected to independently diagnose production incidents across Kubernetes, Linux, networking, containers, storage, and cloud infrastructure, participate in incident response and SLA coverage, and work directly with customers and GPU infrastructure providers when necessary.
We value engineers who actively use modern AI coding agents to improve the speed and quality of their engineering and operational work.
What You’ll Do
- Own and troubleshoot production incidents across GPU-backed Kubernetes environments
- Diagnose issues spanning Kubernetes, Linux, networking, containers, storage, and cloud infrastructure
- Investigate Kubernetes scheduling failures involving resource requests and limits, affinity, taints and tolerations, PriorityClasses, quotas, and related configuration
- Troubleshoot container runtimes including containerd, Docker, and OCI runtimes
- Operate and debug Kubernetes networking, including Services, Ingress, DNS, NetworkPolicy, CNI, and load balancers
- Diagnose TCP/IP, DNS, and TLS issues in production environments
- Troubleshoot Kubernetes storage and CSI issues, including PVCs/PVs, block storage, shared filesystems, and mount failures
- Improve monitoring, alerting, and production observability using tools such as Prometheus and Grafana
- Build diagnostic and operational automation using Bash and Python
- Work directly with customer and infrastructure-provider teams during production incidents
- Determine whether failures originate within Saturn Cloud, Kubernetes, customer infrastructure, networking/storage, or underlying GPU infrastructure
- Produce clear technical evidence for escalation when an issue belongs outside Saturn Cloud’s infrastructure layer
What We’re Looking For
Core Skills
- Hands-on experience using AI coding agents and agentic development tools to accelerate engineering, debugging, automation, and operational workflows
- Deep hands-on experience troubleshooting Kubernetes in production
- Strong Linux systems administration and debugging skills
- Strong knowledge of Kubernetes scheduling and resource management
- Experience with containerd, Docker, or other OCI-compatible container runtimes
- Experience with Helm and Kubernetes deployment/configuration management
- Strong understanding of Kubernetes networking
- Strong TCP/IP, DNS, and TLS troubleshooting skills
- Experience troubleshooting Kubernetes storage and CSI
- Experience with Prometheus, Grafana, or comparable observability tooling
- Bash and Python experience for diagnostics and operational automation
- Experience with at least one major managed Kubernetes environment such as EKS, GKE, AKS, or OKE
GPU Knowledge
You don’t need to be the team’s deepest GPU specialist, but you should be comfortable operating GPU workloads and identifying when an incident has moved beyond the Kubernetes or container layer.
Experience should include:
- NVIDIA GPU Operator and Kubernetes device plugins
- GPU resource allocation and scheduling
- nvidia-smi and basic GPU health diagnostics
- NVIDIA Container Toolkit/runtime
- CUDA and driver compatibility concepts
- Identifying GPU, driver, or hardware failures that require deeper infrastructure escalation
Strong Pluses
- Cilium or Calico
- Terraform
- KAI Scheduler, Grove, Volcano, or other batch/GPU schedulers
- vLLM, NVIDIA Dynamo, Triton, or other inference systems
- Multi-tenant Kubernetes platforms
- Experience operating infrastructure against contractual availability SLAs
- Experience working directly with enterprise customers during production incidents
What Success Looks Like
When an inference endpoint becomes unavailable, you can systematically trace the issue through the load balancer, ingress, Kubernetes Service, workload, scheduler, container runtime, and node.
You can determine whether the root cause belongs to Saturn Cloud, Kubernetes, the customer’s network or storage environment, or the underlying GPU infrastructure, and provide useful evidence to the appropriate team.
You’re comfortable being the primary incident owner, not simply collecting logs and passing the problem to someone else.
Why Saturn Cloud
- Remote-first culture with a high-trust, high-ownership environment
- Work on production infrastructure at the center of modern AI and GPU computing
- Solve technically challenging reliability problems across Kubernetes, cloud, and GPU infrastructure
- Help shape the operational practices of a rapidly evolving AI infrastructure platform
Compensation & Benefits
- Competitive salary
- Flexible PTO
- Fully remote position