Beam logo

GPU Cluster Infrastructure Engineer

Salary
$10.5K–$18K/mo
Hiring from
United States
Work type
Remote
Posted
Is this job info correct?
Show job description

Beam is an ultrafast AI inference platform. We built a serverless runtime that launches GPU-backed containers in less than 1 second and quickly scales out to thousands of GPUs. Developers use our platform to serve apps to millions of users around the globe. We're backed by Y Combinator, Tiger Global, and prominent developer-tool founders, including the founder of Snyk and former CTO of GitHub.

About the Role

We're building out our own GPU capacity and we're looking for an experienced contractor to help us stand up high-performance GPU clusters. The work runs from design review through bring-in, and you'll leave behind the operational foundation our team needs to run them.

  • Review cluster designs and bills of materials across compute, networking, and storage, and catch gaps before hardware is ordered.
  • Lead acceptance testing: validate cabling and optics, bring up the InfiniBand fabric, run burn-in, and hold vendors to their deliverables.
  • Stand up and validate high-performance storage alongside vendor teams.
  • Build the out-of-band management layer and firmware baselines, and secure the management plane for customer-facing environments.
  • Integrate hardware, fabric, and storage telemetry into our observability stack, with alerting and automated health checks.
  • Write runbooks, as-builts, and remote-hands procedures.
  • Provide escalation support after go-live and help our team ramp up.

Skills & Experience

  • You've built and operated NVIDIA HGX or DGX clusters in production at a GPU cloud, HPC center, or AI lab.
  • Hands-on experience with InfiniBand: subnet management and UFM, fabric bring-up, and diagnosing degraded links and optics. NDR or newer.
  • GPU node bring-up and burn-in: firmware, BMC/Redfish, DCGM, NCCL testing, PXE and imaging, and XID error triage.
  • Parallel storage experience: WEKA, VAST, GPFS, Lustre, or similar.
  • Equally effective on the data center floor and remotely, including directing colo remote hands.
  • You troubleshoot methodically across hardware, fabric, and software, document as you go, and communicate clearly with technical and non-technical people.
  • Bonus: recent-generation NVIDIA platforms, bare-metal cloud operations, Ansible or similar automation, Prometheus/Grafana, NVIDIA certifications.

Benefits

  • Competitive salary and meaningful equity
  • Join a fast-growing pre-series A company at the ground floor
  • Health, dental, and vision benefits with 90% coverage for you and 50% for dependents
  • Opportunities to participate in events across the cloud native community
  • Fitness stipend, learning budget, and much, much more

Visa: US citizen/visa only.

Similar jobs

Apply for this job