Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Remotestar Team logo

Senior System Engineer (Munich, Germany)

Remotestar Team
Posted May 27, 2026, 7:35 PM UTC
🇩🇪Germany🏢Hybrid📁Engineering & Development
Is this job info correct?

About client : Well-funded and fast-growing deep-tech company founded in 2019. We are the biggest Quantum Software company in the EU. They are also one of the 100 most promising companies in AI in the world (according to CB Insights, 2023) with 150+ employees and growing, fully multicultural and international. Requirements Systems Programming Expertise: 10+ years of software engineering experience with strong proficiency in Python. You must be comfortable building system agents, APIs, and CLI tools. Deep Kubernetes Knowledge: You understand K8s internals beyond simple deployment. Experience with Custom Resource Definitions (CRDs), Operators, and the Kubernetes API server architecture. GPU Ecosystem Experience: Hands-on experience managing NVIDIA GPU clusters. Familiarity with NVIDIA drivers, CUDA toolkit, and the container runtime (NVIDIA Container Toolkit). Linux Internals: Deep understanding of the Linux kernel, cgroups, namespaces, and system performance tuning. Infrastructure as Code: Mastery of declarative infrastructure tools (Terraform, Ansible) but with a focus on provisioning physical hardware rather than just cloud VMs. Problem Solving: A proven track record of debugging complex distributed systems where the root cause could be code, network, or silicon. Preferred qualifications HPC Background: Experience working with traditional supercomputing schedulers (Slurm, PBS) or modern batch schedulers (Volcano, Kueue, Ray). Bare Metal Provisioning: Experience with tools like Cluster API (CAPI), Metal3, Tinkerbell, Canonical MaaS, or OpenStack Ironic. High-Speed Networking: Knowledge of RDMA, InfiniBand, GPUDirect, and how to expose these technologies to containerized workloads. AI/ML Familiarity: Understanding of how distributed training works (e.g., PyTorch Distributed, Megatron-LM, DeepSpeed) and the infrastructure requirements of Large Language Models (LLMs). Observability: Experience building monitoring for hardware health (DCGM) and distributed tracing for long-running jobs. Location: Applicants must have legal authorization to work in the country where the position is based What you will be doing Building the Control Plane: Designing and developing the software layer (APIs, Controllers, Agents) that automates the lifecycle of bare-metal AI infrastructure. Orchestrating High-Scale Compute: Architecting scheduling solutions for large-scale distributed training jobs across massive clusters of GPUs (NVIDIA H200/B200/B300), ensuring efficient bin-packing and gang scheduling. Optimizing the Fabric: Tuning the software-defined networking layer to support low-latency interconnects (InfiniBand/RDMA/RoCEv2) essential for multi-node training. Developing Kubernetes Extensions: Writing custom Kubernetes Operators and CRDs to abstract complex hardware realities (topology awareness, GPU partitioning) into usable interfaces for our Data Scientists. Hardware-Level Debugging: Investigating and resolving deep systems issues, ranging from PCIe bus errors and NCCL communication timeouts to kernel panics on bare-metal nodes. Defining Standards: Creating the "Golden Image" for AI workloads, managing drivers, firmware, and OS optimizations to squeeze maximum performance out of the hardware. Perks & Benefits Indefinite contract. Equal pay guaranteed. Variable performance bonus. Signing bonus. Relocation package (if applicable). Private health insurance. Eligibility for educational budget according to internal policy. Hybrid opportunity. Flexible working hours. Working in a high paced environment, working on cutting edge technologies. Career plan. Opportunity to learn and teach. Progressive Company. Happy people culture

Similar jobs

Similar jobs

EG

Systemingenieur Engineering (m/w/d)

encontec GmbH

🇩🇪Germany5 hours ago
AG

Masterand (m/w/d) Nachhaltigkeit in Model-Based Systems Engineering

Avelion Group

🇩🇪Germany5 hours ago
Candidate Experience site for interns logo

Presales Systems Engineer - KRITIS Utilities

Candidate Experience site for interns

🇩🇪Germany15 hours ago
Candidate Experience site for interns logo

Presales Systems Engineer - Regional Institutes

Candidate Experience site for interns

🇩🇪Germany15 hours ago
Crowdconsultants logo

Ground Segment System Engineer (IV&V)

Crowdconsultants

🇩🇪Germany17 hours ago
Exoscale logo

Site Reliability Engineer - Compute System & Network (f/m/d)

Exoscale

🌍Germany, Spain18 hours ago