Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
BN

Senior Infrastructure Engineer - GPU Compute

Boundless Networks Inc
Posted 6 hours ago
🇺🇸United States🏠Remote💰$250.0K📁Engineering & Development
Is this job info correct?

Boundless is coordinating GPU compute at scale as it becomes a leader in AI. As a Senior Infrastructure Engineer (GPU Compute), you'll build and operate the compute fabric that powers our AI inference workloads — a large, heterogeneous, globally distributed GPU fleet spanning consumer cards (including RTX 5090) and datacenter hardware. Your job is to keep that fleet full, fast, cheap, and always on: orchestrating workloads across regions and providers, squeezing every bit of performance out of the hardware, and driving down cost per GPU-hour. This role rewards engineers who want to go deep on bare-metal and GPU optimization. You should be comfortable operating with a high degree of autonomy, navigating ambiguity, and defaulting to a strong bias for action. What You'll Do GPU Fleet Orchestration: Operate a heterogeneous, multi-region GPU fleet (consumer + datacenter, including RTX 5090) using tools like SkyPilot, Kubernetes/k3s, and cloud + on-prem providers. Build the patterns that let us schedule inference workloads across the entire fleet reliably. Compute Scheduling & Utilization: Maximize GPU utilization across inference workloads. Own workload placement across spot, on-prem, and cloud capacity, keeping the "always-on inference substrate" saturated and economical. Bare-Metal & GPU Optimization : Go deep on GPU performance — PCIe P2P, ReBAR, NUMA topology (e.g. EPYC SP5), CUDA/driver tuning, memory configuration, and network topology — to push throughput per node. Reliability, Access & Observability : Build secure fleet access (Tailscale, Teleport), robust observability and alerting, and zero-downtime rollouts across a distributed node fleet. Cost Optimization: Drive down $/GPU-hr through spot instance management, intelligent workload placement between on-prem and cloud, and resource scheduling — without sacrificing reliability. 5+ years of infrastructure/DevOps experience operating large-scale production systems Deep expertise in Kubernetes, Docker, and container orchestration at scale Strong Linux systems administration skills Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi) Track record of managing mission-critical, high-throughput systems Strong infrastructure-as-code background in heterogeneous environments Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.) Comfort navigating ambiguity with a strong bias for action Nice to Have Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning) Experience operating ML training or other large-scale distributed compute infrastructure Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm) Familiarity with fleet access and networking tooling (Tailscale, Teleport) Knowledge of network optimization and topology design Experience with multi-region, globally distributed systems Proficiency in Rust or low-level systems programming Experience with on-premises data center operations Additional Requirements Candidates must include a public GitHub profile in their application. The GitHub profile should demonstrate a minimum of 1 year of activity/history. Applications that do not include a GitHub profile, or show insufficient activity, will not be considered. At Boundless, we take care of our people, because building the future of AI compute starts with an empowered team. Here's what you can expect when you join us: Competitive salary (proposed band b/t US$175k and $250k annually) + equity allocation Health, dental, vision (for U.S. employees; region-adjusted globally) Flexible PTO Professional development and conference travel budget Remote-first with regular off-sites and a high-trust, high-velocity team environment We are a global team, and applicants from around the world are welcome to apply.

Similar jobs

Similar jobs

Roblox logo

Principal Software Engineer, GPU Compute

Roblox

🇺🇸United StatesJun 4, 2026, 6:51 PM UTC
Nebius logo

Senior Systems Software Engineer, GPU Compute

Nebius

🇺🇸United StatesMay 28, 2026, 12:34 AM UTC
Lightning AI logo

Infrastructure Engineer (GPU & Compute)

Lightning AI

🌍United Kingdom, United StatesMay 27, 2026, 7:41 PM UTC
Agileengine logo

QA/Support Engineer ID80988

Agileengine

🇺🇸United States1 hour ago
Agileengine logo

Senior AppSec Engineer ID71672

Agileengine

🌍Brazil, Mexico, United States1 hour ago
Main Line Talent Group logo

Infor M3 Developer

Main Line Talent Group

🌍Canada, United States2 hours ago