Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Reactor logo

Member of Technical Staff, DevOps

Reactor
Posted 3 hours ago
🛂Visa sponsorship
🇺🇸United States
📁Engineering & Development
Is this job info correct?

We're looking for a DevOps / SRE engineer to own the reliability, delivery, and observability of our AI platform. You'll be the person who ensures models get from a developer's branch to production without anyone losing sleep — and when something does go wrong at 2am, you'll be the one who knows where to look. We run production across multiple Kubernetes clusters, cloud providers, and regions. Our deployment pipeline is fully automated through CI/CD and GitOps, our infrastructure is managed as code, and our observability stack gives us full visibility across every service and GPU workload. This role is about making all of that faster, more reliable, and easier to operate as we scale. What You'll Do Own and evolve our CI/CD pipelines: dynamic pipeline generation across a monorepo of Go services, Python model containers, and Helm charts Operate and improve our GitOps deployment lifecycle: Helm releases, Kustomizations, and image automation across multiple clusters Build and maintain our observability stack: distributed tracing, metrics, dashboards, and alerting across all services and GPU workloads Define and track SLOs for core platform services, including session latency, model cold start time, and streaming reliability Run incident response: triage production issues, write postmortems, build runbooks, and drive reliability improvements Manage infrastructure-as-code across multiple cloud providers and regions: plan/apply workflows, state management, drift detection Operate secret management: encrypted secrets, external secret syncing, certificate automation Improve deployment safety: canary rollouts, health checks, startup probes, rollback automation Manage authentication infrastructure: OIDC federation for CI, workload identity for cloud services, cross-cloud credential management Participate in on-call rotation and build the tooling that makes on-call less painful What We're Looking For You've run production Kubernetes clusters and been on-call for them. You've debugged node scheduling failures, OOM kills, and mysterious pod evictions at 3am Strong CI/CD experience: you've built and maintained pipelines for monorepos, not just single-service repos GitOps experience: you understand reconciliation loops, drift detection, and why image automation matters Infrastructure-as-code fluency with Terraform or similar across multiple environments and cloud accounts You know observability beyond just "set up dashboards". You've defined SLOs, built alerting that doesn't page on noise, and used traces to debug cross-service latency issues. Comfortable with secret management patterns (KMS, encrypted configs, external secret operators). You've thought about credential rotation and zero-trust. Incident response experience: you've triaged production outages, written postmortems that actually led to improvements, and built runbooks that other engineers could follow You write code, not just YAML. Proficiency in Go, Python, or Bash for building tooling, automation, and pipeline scripts Nice to Have Experience with GPU workloads on Kubernetes: device plugins, GPU-aware scheduling, GPU monitoring Multi-cloud operations beyond a single provider Real-time or streaming workloads: low-latency systems where p99 matters more than average Experience with Helm chart authoring and managing complex value layering across environments Familiarity with real-time media or relay infrastructure FinOps experience: GPU cost optimization, spot/preemptible instance management What We're Not Looking For Engineers who treat infrastructure-as-code as "click around in the console and import later" SREs who've only monitored systems but never built the deployment pipelines that ship to them Candidates whose CI/CD experience is limited to GitHub Actions for a single-service repo People who write alerts that fire every day and then get ignored Logistics We are based in-person in San Francisco. We are also hiring for this role in Europe for on-call coverage and timezone distribution. Benefits Competitive salary and meaningful early equity Visa sponsorship and relocation support Generous health, dental, and vision coverage

Similar jobs

Similar jobs

Capitalone logo

Senior Lead Software Engineer, iOS

Capitalone

🇺🇸United States43 minutes ago
Capitalone logo

Senior Manager, Software Engineering (Java, Golang, AWS)

Capitalone

🇺🇸United States43 minutes ago
Medtronic logo

Senior Quality Engineer

Medtronic

🌍Singapore, United States1 hour ago
Medtronic logo

Principal IT Technical Architect - SAP S/4 HANA Implementation Team

Medtronic

🇺🇸United States1 hour ago
AECOM logo

Structural Engineer (Substation Focus) - Remote (US)

AECOM

🇺🇸United States3 hours ago
AECOM logo

Senior Structural Engineer (Substation Focus) - Remote (US)

AECOM

🇺🇸United States3 hours ago