Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
5Cai logo

Engineering Manager, GPU Infrastructure

5Cai
Posted 2 days ago
🇺🇸United States🏠Remote📁Engineering & Development
Is this job info correct?

Industry: Hyperscale and AI Data Center and Cloud Computing Location: Remote (US, Pacific Time Zone) Employment Type: Full-Time Reporting to: VP, Operations POSITION SUMMARY We are seeking an experienced Engineering Manager, GPU Infrastructure to lead the planning, deployment, integration, and operational readiness of large-scale AI infrastructure environments. This role is responsible for delivering production-grade GPU clusters that support AI training, inference, and high-performance computing workloads across cloud, hybrid, and on-premises environments. The ideal candidate brings deep technical expertise in GPU infrastructure, networking, storage, automation, and datacenter deployment, combined with strong program leadership and cross-functional execution skills. This leader will oversee end-to-end AI cluster deployment initiatives, including hardware integration, rack-and-stack operations, provisioning automation, performance validation, and operational handoff. The role requires hands-on familiarity with modern AI infrastructure tooling and architectures, including Canonical MaaS, VAST Data storage platforms, and both InfiniBand and Ethernet-based GPU networking fabrics. KEY RESPONSIBILITIES AI and GPU Cluster Deployment & Delivery · Oversee and partake in deployment and integration of GPU-based compute platforms from NVIDIA and other accelerator vendors · Lead and participate in end-to-end logical deployment of large-scale AI and GPU clusters in state of the art datacenters. · Manage deployment programs spanning compute, storage, networking, power, cooling, and automation layers. · Participate in cluster architecture review for AI training, inference and distributed compute workloads · Coordinate rack-and-stack and cabling sequencing, network deployment, burn-in testing, and cluster validation.Validate deployment readiness, topology consistency, GPU fabric performance, acceptance testing, and operational turnover processes. · Establish repeatable and documented deployment methodologies and scalable operational standards. Networking & Fabric Management · Lead deployment and operational validation of high-performance GPU interconnects using InfiniBand and Ethernet GPU fabric architectures · Ensure proper implementation of: spile-leaf architectures, RDMA, network telemetry and performance tuning · Coordinate closely with network engineering teams on topology implementation and performance optimization. Storage & Data Infrastructure · Coordinate with storage engineering teams on deployment and integration of high-performance storage environments supporting AI workloads. · Ensure successful implementation and operational optimization of data storage platforms · Validate storage throughput, latency, and GPU data delivery performance. Automation & Provisioning · Lead infrastructure automation initiatives for cluster provisioning and lifecycle management. · Manage deployment tooling and orchestration platforms including: o Infrastructure-as-Code frameworks o Automated imaging and provisioning systems (e.g. Canonical MaaS) o Cluster monitoring and observability tools · Drive standardization and deployment automation to improve speed, reliability, and repeatability. Leadership & Program Management · Build and lead high-performing technical deployment and infrastructure engineering teams. · Partner with datacenter operations, hardware vendors, networking teams, and AI platform engineering groups. · Establish strong Project Management Office (PMO) partnership while driving consistent, accurate project updates across the team and systems (e.g. Jira) · Develop operational procedures, documentation, and deployment best practices. · Mentor engineers and technical leads across infrastructure domains. QUALIFICATIONS REQUIRED · Bachelor's degree in Computer Science, Engineering, Information Technology, or related field (or equivalent experience). · 10+ years of infrastructure engineering or datacenter deployment experience. · 5+ years leading deployment or operations teams supporting large-scale AI, HPC, or GPU infrastructure. · Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments. · Strong expertise with: o Canonical MaaS o Data storage platforms o InfiniBand and Ethernet GPU fabrics o Network architecture o Linux systems administration o GPU server architectures · Strong understanding of: o RDMA and RoCE networking o High-performance storage architectures o Cluster automation and provisioning o Datacenter infrastructure operations · Proven ability to manage complex cross-functional infrastructure deployment programs. Preferred Qualifications · Experience deploying NVIDIA DGX SuperPOD or similar AI infrastructure solutions. · Familiarity with: o NVIDIA networking technologies o Spectrum-X or Quantum platforms o AI model training infrastructure o Liquid cooling environments o DCIM and observability platforms · Experience in hyperscale, cloud, or AI infrastructure environments. · Certifications in networking, Linux, Kubernetes, or cloud infrastructure are a plus. Key Competencies · Technical leadership · Infrastructure architecture · Program execution · Cross-functional collaboration · Vendor and stakeholder management · Problem-solving under operational pressure · Process improvement and automation · Excellent communication and documentation skills 5C Data Centers is an equal opportunity employer. We evaluate all qualified applicants without regard to race, religion, gender, age, national origin, disability, sexual orientation, veteran status, or other protected status. #LI-LS1

Similar jobs

Similar jobs

Liquid Ai logo

Member of Technical Staff - GPU Infrastructure Engineer

Liquid Ai

🇺🇸United States3 days ago
Roblox logo

Senior Hardware Engineer - GPU & AI Infrastructure

Roblox

🇺🇸United StatesJun 18, 2026, 9:05 AM UTC
Cohere logo

Software Engineer, GPU Infrastructure (HPC)

Cohere

🌍Canada, United StatesMay 27, 2026, 10:08 PM UTC
Primeintellect logo

Member of Technical Staff - GPU Infrastructure

Primeintellect

🇺🇸United StatesMay 27, 2026, 8:50 PM UTC
deCircle logo

Hyperbolic Labs - Senior GPU Infrastructure Engineer

deCircle

🇺🇸United StatesMay 27, 2026, 8:44 PM UTC
CoreWeave logo

Sr GPU Infrastructure Software Engineer

CoreWeave

🇺🇸United StatesMay 27, 2026, 7:45 PM UTC