Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
iFrame logo

Cluster Site Reliability Engineer

iFrame
Posted 3 hours ago
🛂Visa sponsorship
🇨🇦Canada
💰$210.0K–$340.0K📁Engineering & Development
Is this job info correct?
Cluster & SRE

Cluster Site Reliability Engineer

Own the physical reality of the platform in one of our seven regions. You bring up new GPU racks, validate InfiniBand fabric end-to-end, and keep the cluster running at the SLA. This is a hands-on role; you will see the hardware.

Apply by email All open roles

Team Cluster & SRE

Location On-site

  • Cologix region of choice (Toronto, Montréal, Columbus, Vancouver, Ashburn)

Type Full-time

  • IC4–IC6

Stack LinuxInfiniBand (NDR / XDR)NCCL / RCCLKubernetes (host-level)TerraformPrometheus / GrafanaGo or Python

Apply by email

Hiring manager replies within 5 business days.

The team

About The Team

Cluster SRE is six engineers across the seven regions. Each region has a primary and a secondary; you will be one of those for your region. The team coordinates daily, deploys weekly, and rotates a global pager.

Reports to the head of cluster engineering. Primary on a single region; rotates secondary for one neighboring region.

The role

What you'll do

  • 01 Bring up new B200 / B300 / MI300X racks: cabling, ToR config, NCCL/RCCL all-reduce validation, MFU baseline tests.
  • 02 Drive InfiniBand fabric to spec — NDR / XDR depending on the rack — and chase residual bit-error budget down to zero.
  • 03 Run capacity planning across seven regions: forecast demand, model power and thermal headroom, work with procurement on lead times.
  • 04 Own the regional incident response. P1 incidents page within fifteen minutes; resolution target is four hours.
  • 05 Build and maintain the bring-up runbook so the second hire after you can do their first rack solo.
  • 06 Carry the global pager about one week per six, alongside runtime and customer engineering.

The bar

What we're looking for

  • Five-plus years operating large compute clusters — supercomputing centers, hyperscaler infra, or HPC at a national lab count.
  • Deep InfiniBand and Ethernet RoCE experience: subnet manager tuning, fabric debugging, lossless networking.
  • Strong Linux fundamentals: kernel parameters, NUMA topology, kernel cgroup limits.
  • Comfort writing Go or Python for tooling. We are not strict about which.
  • Calm under load. You will be the named person on a $50M-ARR account when something goes wrong.

Bonus

Nice to have, not required

  • DGX H100 / H200 / B200 bring-up history.
  • Experience with NVIDIA Bright / Base Command Manager.
  • Bare-metal provisioning systems: Tinkerbell, MAAS, Razor, or in-house equivalents.
  • Procurement / vendor management experience.

Compensation

In writing, like everything else

We publish bands. We meet them. The number you see on the offer is the same number your future peers got at the same level. We do not negotiate; we level.

Base

$210,000 – $340,000 USD (US Cologix regions) / equivalent in CA.

Equity

Meaningful early-stage equity, refreshed on tenure milestones.

Notes

On-site pay differential at Cologix regions outside SF / NYC / Bay Area is +5–10% to compensate for travel.

How to apply

One email is enough

Send a short note to [email protected] with the role title in the subject line. Include your CV or LinkedIn, one or two links to work you're proud of, and a sentence on why this role specifically. Hiring managers reply within five business days, regardless of outcome.

  • 01

Application

A hiring manager reads every email. Reply within five business days.

  • 02

Manager call

30–45 minutes. Scope, role, mutual fit. We share the comp band on this call.

  • 03

Technical loop

3–4 sessions on the same day. Real problems, no homework, no whiteboard riddles.

  • 04

Offer

Same-week offer at the published band for your level. Start dates are flexible.

Equal opportunity

We hire on the work. Race, gender, age, nationality, religion, sexual orientation, disability, and veteran status do not factor into our decisions. We sponsor visas for senior roles in the US, UK, and EU — bring it up on the manager call.

Need an accommodation for the interview process? Mention it in your application or write [email protected].

Also open

Other roles you might consider

  • Runtime

Inference Acceleration Engineer

Ship the next 2× on the open-source model catalog. Triton, CUDA, ROCm. You will publish what you ship.

View role

  • Research lab

Distributed Training Researcher

Lead a paper / quarter on multi-thousand-GPU pre-training. Co-appointment with a partner university available.

View role

  • Customer engineering

Customer Cluster Engineer

Embedded with a small portfolio of reserved-tier accounts. Distributed training perf, NCCL, kernel tuning.

View role

All open roles

One last thing

If this role isn't quite right but you'd be a fit at iframe.ai, write anyway.

Senior engineers and researchers can apply outside the listed roles. The bar is the same. The reply window is the same.

Apply: Cluster Site Reliability Engineer General application

Similar jobs

Similar jobs

Capitalone logo

Distinguished Engineer

Capitalone

🌍Canada, India, Mexico, United States3 hours ago
AP

IDPS Cyber Automation Analyst

Applyglobal

🇨🇦Canada4 hours ago
Diligent Corporation logo

Software Engineer AI - Multiple Levels

Diligent Corporation

🇨🇦Canada4 hours ago
lululemon logo

Technical Developer - Women's Lifestyle

lululemon

🇨🇦Canada2 days ago
SC

Flotilla Engineer / Mate

Seafarer Cruising and Sailing Holidays

🇨🇦Canada2 days ago
lululemon logo

Senior Software Engineer - DevOps, Digital Commerce Services

lululemon

🇨🇦Canada4 days ago