5Cai logo

Senior AI GPU Deployment Engineer

Salary
$120K–$150K
Hiring from
United States
Work type
Remote
Posted
Is this job info correct?
Show job description

5C DATA CENTERS

Join the Future of Digital Infrastructure

Are you a passionate Senior AI GPU Infrastructure Engineer looking to make a meaningful impact? We're building the next generation of digital infrastructure powering hyperscalers, AI innovation, and high-performance computing across North America.

JOB TITLE

Senior AI GPU Infrastructure Engineer

DEPARTMENT

Cloud Operations

LOCATION

USA

WORK ARRANGEMENT

Remote

TRAVEL REQUIREMENTS

Minimal (less than 10%)

SALARY RANGE

$120,000 – $150,000

Role Summary

We are seeking an experienced Senior AI GPU Infrastructure Engineer to plan, deploy, and operationalize large-scale GPU AI infrastructure environments. This role delivers production-grade GPU clusters supporting AI training, inference, and high-performance computing workloads in our hyperscale data centers.

The ideal candidate brings deep technical expertise in GPU infrastructure, network fabrics, storage, automation, and Linux systems administration, combined with strong execution and troubleshooting skills.

How We Work at 5C

Our core values guide how we collaborate, make decisions, support one another, and serve our customers. We're looking for people who embrace them and help us build something great.

What You Will Do

AI and GPU Cluster Deployment

  • Deploy and integrate GPU-based compute platforms from NVIDIA and other accelerator vendors
  • Execute end-to-end deployment of multi-rack AI clusters in hyperscale datacenters
  • Support rack-and-stack, NVLink/NVSwitch cabling, fabric deployment, burn-in, and cluster validation
  • Validate deployment readiness, GPU fabric performance, acceptance testing, and operational handoff

Networking and Fabric Management

  • Validate high-performance GPU interconnects based on InfiniBand and Ethernet GPU fabric architectures
  • Deploy fabric configuration engines (Subnet Manager), observability platforms (UFM) and validate fabric performance (nccl)
  • Collaborate with network engineering teams on topology implementation and optimization

Storage and Data Infrastructure

  • Coordinate with storage engineering teams on deployment and integration of high-performance storage environments supporting AI workloads (e.g. VAST Data)
  • Validate storage throughput, latency, and GPU data delivery performance

Firmware and Systems Configuration

  • Manage firmware updates for GPUs, NICs, BMC, and other components across large-scale clusters
  • Configure and validate BIOS settings for HPC/AI workloads (CPU affinity, NUMA, C-states, power management)
  • Configure and manage BMC for remote infrastructure management (IPMI, Redfish)

Automation and Provisioning

  • Contribute to infrastructure-as-code automation development for cluster provisioning and lifecycle management
  • Contribute to improving and documenting repeatable deployment methodologies and scalable operational standards
  • Query and analyze deployment outcomes using SQL for diagnostics and operational reporting

What You Bring

  • Bachelor's degree in Computer Science, Engineering, IT, or related field (or equivalent experience)
  • 5+ years of infrastructure engineering or datacenter deployment experience
  • 3+ years deploying large-scale AI, HPC, or GPU infrastructure
  • Hands-on experience deploying and operating large GPU clusters in enterprise or hyperscale environments
  • Strong expertise with:
    • GPU architectures
    • InfiniBand (NDR/XDR) and Ethernet GPU fabrics (Spectrum-X)
    • NVLink, NVSwitch, and GPU-direct technologies
    • Canonical MaaS and automated provisioning systems
    • VAST Data or similar high-performance storage platforms
    • Linux systems administration for HPC/AI workloads
    • Infrastructure-as-Code and configuration management (Ansible)
    • Python, Shell, and SQL for infrastructure automation and diagnostics
  • Strong understanding of:
    • RDMA, RoCE, and lossless Ethernet fabrics
    • Cluster automation, observability, and lifecycle management

Why Join 5C Data Centers?

At 5C, we believe great people build great companies. You'll build a rewarding career while helping shape the future of digital infrastructure - one of the fastest-growing industries in the world.

Career Growth

Build a rewarding career in one of the world's fastest-growing industries.

Industry Leadership

Help power the infrastructure behind AI and high-performance computing.

Entrepreneurial Culture

Your ideas matter - we empower employees to help shape our future.

Comprehensive Rewards

Competitive pay plus meaningful, lasting impact on the work you do.

Life at 5C

We're more than a workplace - we're a team of builders, innovators, and problem-solvers united by a shared purpose: creating infrastructure that powers the technologies transforming our world. Your voice matters here, and we encourage fresh ideas at every level.

Ready to Apply?

If this opportunity sounds like the right fit for you, we'd love to hear your story. Apply today and discover what your future could look like at 5C Data Centers.

5C Data Centers is an equal opportunity employer.

We celebrate diversity and are committed to creating an inclusive environment where everyone can thrive. 5C evaluates qualified applicants without regard to race, color, religion, gender, national origin, age, sexual orientation, gender identity or expression, disability status, or any other legally protected characteristic.

Similar jobs

Apply for this job