Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Boson AI logo

Site Reliability Engineer

Boson AI
Posted 2 weeks ago
🇨🇦Canada🏠Remote💰CA$125.0K–CA$250.0K📁Engineering & Development
Is this job info correct?

About The Role Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work. Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from “it works” to dependable, observable, and scalable. You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area—networking, cluster scheduling, storage, GPU systems, or AI infrastructure— and the curiosity and judgment to collaborate across the rest. Responsibilities Design, operate, and improve reliable infrastructure for AI training and inference workloads Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements Improve provisioning, configuration management, testing, and deployment automation Help plan cluster growth, capacity allocation, upgrades, and lifecycle management Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards Minimum Qualifications 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role Strong hands-on expertise in at least one of the following: Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms Distributed storage, particularly Ceph GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting AI training or model-serving infrastructure Experience operating production systems with a focus on availability, performance, security, and automation Strong Linux administration and scripting skills A systematic approach to troubleshooting across multiple layers of a complex system Clear written and verbal communication skills, including the ability to work effectively with a distributed team Preferred Qualifications Experience supporting GPU-intensive AI or HPC environments Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling Experience operating or tuning Ceph clusters Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems Experience with hardware provisioning, firmware management, and bare-metal automation Experience running large-scale distributed training or high-throughput inference workloads Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we’d love to hear from you.

Similar jobs

Similar jobs

MO

Senior Site Reliability Engineer

Morningstar

🇨🇦Canada3 hours ago
Coinbase logo

Senior Software Engineer, Core Reliability

Coinbase

🇨🇦CanadaYesterday
Careers2 Chemtradelogistics logo

Senior Mechanical Reliability Engineer

Careers2 Chemtradelogistics

🌍Canada, United StatesYesterday
ClickHouse logo

Senior Site Reliability Engineer- Remote

ClickHouse

🌍Australia, Canada, Singapore, United Kingdom, United States4 days ago
Lumenalta logo

Senior DevSecOps / Site Reliability Engineer (GCP)

Lumenalta

🌍Canada, Colombia, Dominican Republic, Mexico5 days ago
Lightspeedhq logo

Staff Site Reliability Engineer

Lightspeedhq

🇨🇦Canada1 weeks ago