Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Andromeda logo

Customer Reliability Engineer

Andromeda
Posted May 28, 2026, 3:42 AM UTC
🌍Worldwide🏠Remote📁Engineering & Development
Is this job info correct?

Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco · Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers. We began with a single managed cluster — but it filled almost instantly. Since then, we’ve been quietly building the systems, network, and orchestration layer that makes the world’s AI infrastructure more accessible. Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it’s needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth. Our long-term vision is to build the liquidity layer for global AI compute — a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets. We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering. What You’ll Do Provision, configure, and operate Kubernetes-based clusters for customers across multiple providers. Build automation and tooling to streamline cluster deployments and integrations. Debug customer issues across networking, storage, scheduling, and system layers. Improve reliability and scalability of both training and inference infrastructure. Design and implement monitoring, alerting, and observability for critical systems. Collaborate with engineering and product teams to plan and deliver infrastructure for new services. Participate in on-call and incident response, leading postmortems and reliability improvements. What We’re Looking For 5+ years experience in SRE, DevOps, or infrastructure engineering roles. Strong Linux systems and networking fundamentals. Deep experience with Kuber Kubernetes and container orchestration at scale. Proficiency with Infrastructure-as-Code (Terraform, Helm, Ansible, etc.). Strong automation and scripting skills (Python, Go, or Bash). Experience with observability stacks (Prometheus, Grafana, Loki, Datadog, etc.). Track record of operating production systems and leading incident response. Nice to Have Exposure to ML/AI infrastructure or GPU-based systems (CUDA, Slurm, Triton, etc.). Familiarity with high-performance networking (InfiniBand, NVLink) or distributed storage (VAST, Weka, Ceph). Customer-facing support or consulting experience. Why You’ll Love It Here This is a builder’s role. You’ll have ownership and autonomy to shape how our systems run, working directly with customers and providers while building the foundation for reliable, scalable AI infrastructure.

Similar jobs

Similar jobs

Miris logo

Site Reliability Engineer (Senior+)

Miris

🌍Worldwide2 weeks ago
Cloudlinux logo

Senior Database Reliability Engineer (DBRE) (worldwide remote)

Cloudlinux

🌍Worldwide2 weeks ago
Chainlink Labs logo

Senior Site Reliability Engineer, DevEx

Chainlink Labs

🌍Worldwide2 weeks ago
株式

fixed-term project engagement|Site Reliability Engineer (SRE)

株式会社天地人

🌍Worldwide2 weeks ago
Chainlink Labs logo

Senior Site Reliability Engineer, CCIP

Chainlink Labs

🌍WorldwideJun 30, 2026, 2:55 AM UTC
Luuplicareers logo

Site Reliability Platform Engineer (SRE)

Luuplicareers

🌍WorldwideJun 28, 2026, 3:08 PM UTC