Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Verda logo

HPC Engineer

Verda
Posted 8 hours ago
🇩🇪Germany🏠Remote📁Engineering & Development
Is this job info correct?

HPC Engineer At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work. We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco. Join Verda while it’s still being built - not once it’s finished. Why Verda Cash and equity compensation along with various fringe benefits (healthcare, lunch, wellbeing, and more). Profitable operations with rapid, sustained growth. 30+ nationalities, with 6 different ones on the management team. A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem. Practicalities Work mode: Remote (EU) Level: Mid / Senior Employment type: Full time and permanent About the role GPUs only deliver value once they're wired into a cluster that researchers can actually train on. As our HPC Engineer, you'll own the baremetal and virtualized clusters behind our AI cloud - from the InfiniBand fabric and shared filesystems up through the workload orchestration layer. You'll be the person who keeps large-scale GPU clusters healthy, performant, and ready for the next workload. Your responsibilities Administer baremetal GPU/HPC clusters end to end, from provisioning through day-two operations Administer virtualized clusters, including the hypervisor and GPU virtualization stack underneath them Design, deploy, and tune InfiniBand fabrics, including topology planning, subnet management, and performance validation Deploy and operate shared/parallel filesystems supporting training and inference workloads, balancing performance, capacity, and reliability Troubleshoot and resolve issues across the full fabric stack: fibers, transceivers, NICs, switches, drivers, and firmware Partner with remote-hands and data center teams to diagnose hardware faults and execute physical-layer fixes and cluster expansions Operate and tune Slurm (or equivalent) workload scheduling deployments used by customers and internal teams Keep issue tracking, IPAM, and DCIM records accurate as clusters are built, changed, and decommissioned Participate in on-call rotations and incident response for cluster-level issues Collaborate with platform, network, and storage teams to integrate new clusters into the broader AI cloud Your key competencies Solid Linux skills, with specialization in memory management, PCIE topologies and virtualization being a bonus Deep Infiniband knowledge, including fabric design, subnet management, and performance tuning Solid experience with IB clustering, troubleshooting/debugging, understanding the ecosystem of fibers+transceivers+NICs+switches and all the things that could possibly fail in them Experience about shared filesystems (e.g. Lustre, GPFS/Spectrum Scale, WekaFS, or similar) Ability to work with remote hands teams to diagnose and resolve hardware issues remotely Knowledge about NCCL, CUDA, DOCA and the Nvidia stack Knowledge about Slurm and/or other workload scheduling solutions, bonus points for Slinky/slurm-bridge Understanding the importance of keeping issue tracking/IPAM/DCIM up to date Comfort operating production clusters where uptime and performance directly affect customer workloads Scripting/automation ability (e.g. Python, Bash, Ansible) for repeatable cluster operations Nice to have RoCEv2 knowledge (and/or Spectrum-X) Understanding of agentic guardrails, especially when applied to administration of complex systems Ability to think beyond what is needed right now vs. some given trajectory or roadmap Experience with GPU health-checking and diagnostics tooling (e.g. DCGM, field diagnostics) Experience with baremetal provisioning/orchestration tooling (e.g. MAAS, Foreman, custom PXE/iPXE pipelines) Familiarity with GPU-aware virtualization or containerization (Kubernetes device plugins, KVM/QEMU with GPU passthrough, SR-IOV) Exposure to observability stacks (Prometheus, Grafana, Loki) for cluster-level monitoring What's next We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now. Please submit your application through our Careers page. We don't accept applications sent by email. Department Research & Development Locations EU Remote status Fully Remote About Verda Verda (formerly DataCrunch) is a technology company building the next generation of cloud infrastructure for AI – compute that's instant, on-demand and at scale. Headquartered in Helsinki, the company operates globally across Europe, the US and Asia. Verda employs over 100 people from nearly 30 nationalities and has raised over $200M in total funding from investors including Lifeline Ventures, byFounders, J12 Ventures, Skaala, Varma and Tesi, alongside leading financial institutions.

Similar jobs

Similar jobs

indivHR | We 💚 IT Recruiting logo

Senior HPC Engineer / HPC Architect (m/w/d)

indivHR | We 💚 IT Recruiting

🇩🇪GermanyMay 28, 2026, 3:43 AM UTC
Nebius logo

Senior HPC Engineer, GPU Compute

Nebius

🌍Germany, Netherlands, United KingdomMay 28, 2026, 12:34 AM UTC
OG

Teamleader IT Applications (m/w/d)

ODW-ELEKTRIK GmbH

🇩🇪Germany1 hour ago
Zeissgroup logo

Internship in Next-Generation System Interaction with Eye-Tracking (f/m/x)

Zeissgroup

🇩🇪Germany7 hours ago
VE

Director Clinical Strategy

Vertanical

🇩🇪Germany59 minutes ago
NielsenIQ logo

Senior Product Manager

NielsenIQ

🇩🇪Germany1 hour ago