Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education & Training jobs
  • Remote Healthcare & Nursing jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact mahmoud@relomote.com · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Reactor logo

Platform Engineer

Reactor
Posted 5 hours ago
📦Relocation support🛂Visa sponsorship
🇺🇸United States
📁Engineering & Development
Is this job info correct?

You'll own the infrastructure platform that our AI models run on. This isn't a CI/CD-focused DevOps role. You'll work across GPU orchestration, multi-cloud Kubernetes, real-time networking, and observability. You'll be the person who knows why a model pod took 4 minutes to schedule, why cross-region latency spiked, or why a media relay is dropping packets. We run production today across multiple Kubernetes clusters, regions, and GPU types, and we're actively expanding to additional cloud providers. You'll lead that expansion and keep everything running. What You'll Do Provision and manage multi-region Kubernetes clusters across AWS and GPU cloud providers using infrastructure-as-code. Own the GitOps deployment lifecycle (Helm charts, Kustomize overlays, image automation, and continuous delivery.) Manage GPU node infrastructure: scheduling, model weight caching, image prefetching for fast cold starts, and GPU observability. Operate and improve our networking layer: ingress and gateway management, load balancing, media relay infrastructure, and cross-region connectivity. Build and maintain our observability stack: metrics, logs, traces, and profiling across all services and GPU workloads. Maintain infrastructure security: IAM, secret management, certificate automation, and encryption at rest. Own CI/CD pipelines for monorepo builds spanning Go services, Python model containers, and Helm chart releases. Partner with ML engineers on model serving: container optimization, health checks and startup tuning, media pipeline performance, and multi-GPU configuration. What We're Looking For You've operated Kubernetes in production at scale, not just deployed to it, but debugged node-level scheduling issues, tuned autoscalers, and managed cluster upgrades. Strong infrastructure-as-code experience (Terraform, Pulumi, or similar) across multiple environments and regions. You've worked with GPU workloads on Kubernetes: device plugins, node taints/tolerations, GPU-aware scheduling. You understand why bin-packing matters for expensive hardware. Experience with GitOps tooling (FluxCD, ArgoCD, or similar) and Helm chart authoring. Comfortable with Redis or similar in-memory data stores (replication, persistence, pub/sub or streaming patterns) Familiarity with modern observability stacks (Prometheus, Grafana, OpenTelemetry, or equivalent) and knowing when to reach for metrics vs. logs vs. traces. Solid networking fundamentals: load balancers, TLS, DNS, NAT. Real-time or low-latency networking experience is a strong plus. You've worked in a startup where you owned infrastructure end-to-end, not just one slice of it. Nice to Have Experience with GPU cloud providers beyond AWS (Crusoe, CoreWeave, Lambda Labs, Nebius) Real-time media or streaming infrastructure Go or Python proficiency Familiarity with ML model serving (container image optimization, weight loading, GPU driver and runtime management) FinOps and GPU cost optimization What We're Not Looking For Pure CI/CD pipeline engineers who haven't operated Kubernetes clusters directly Candidates whose infrastructure experience is limited to managed PaaS (Heroku, Vercel, Railway) People who need a fully defined scope, this role requires figuring out what to build next, not just executing tickets Benefits Competitive San Francisco salary and meaningful equity We sponsor visas and support relocation to the US Generous health, dental, and vision coverage

Similar jobs

Similar jobs

Cat logo

Lead Data Engineer – Physical AI Platform, Data Engineering

Cat

🇺🇸United StatesYesterday
Pluralis Research logo

Machine Learning Engineer - ML Training Platform

Pluralis Research

🌍Australia, United States2 days ago
Thinking Machines Lab logo

Software Engineer, Evaluation Platform / Infra

Thinking Machines Lab

🇺🇸United States1 weeks ago
Raydar logo

Senior Healthcare Platform Engineer

Raydar

🇺🇸United States1 weeks ago
Applied Compute logo

AI Platform Engineer

Applied Compute

🇺🇸United States4 weeks ago
Wispr Flow logo

Platform Engineer, Billing Systems

Wispr Flow

🇺🇸United StatesJul 28, 2026, 6:47 AM UTC