Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Madiffpl logo

Site Reliability Engineer (AI)

Madiffpl
Posted May 28, 2026, 10:47 PM UTC
🌍Probably Worldwide🏠Remote📁Engineering & Development
Is this job info correct?

Job Description This is a remote position. We are looking for a Senior Site Reliability Engineer to support advanced AI platforms responsible for production-grade applications and pipelines. The role focuses on building and maintaining reliability, scalability, and operational excellence across multiple AI-driven systems. The engineer will work on a central operational layer for monitoring and managing AI workloads, improving system stability, and reducing incidents. This is a hands-on role requiring direct involvement in diagnosing production issues, implementing fixes, and optimising monitoring, alerting, and CI/CD processes. The position requires close collaboration with engineering teams to improve release quality, standardise telemetry, and ensure stable and predictable system behaviour in a distributed cloud environment. Responsibilities ​ • Build and maintain central monitoring and alerting layer for AI applications and pipelines • Define and implement SLIs, alerts, and operational dashboards • Manage incidents including triage, coordination, root cause analysis, and prevention • Standardise telemetry across systems including latency, throughput, and failures • Optimise CI CD pipelines and introduce quality gates for reliability • Work closely with engineering teams to reduce recurring issues and improve stability Requirements • Minimum 5+ years of experience in SRE, Platform, or Production Engineering • Strong hands on experience with Kubernetes and production environments • Experience with Azure and Azure DevOps • Experience with monitoring tools such as Datadog • Strong understanding of incident management and root cause analysis • Ability to build practical monitoring and alerting systems Nice to have • Experience with AI or LLM pipelines • Experience building monitoring platforms across multiple systems • Experience with Grafana • Experience working in large scale or distributed environments Expectations • Strong ownership mindset and accountability for system stability • Proactive approach to identifying risks and improvements • Hands on engineer actively working with systems, not only coordinating • Comfortable working in dynamic and evolving environments Benefits • Solid, competitive salary • Work in a multinational environment on international projects • Comprehensive healthcare • Long-term B2B contract with a stable project pipeline • Work model: fully remote

Similar jobs

Similar jobs

pod network logo

Site Reliability Engineer (APAC)

pod network

🌍Probably WorldwideJun 19, 2026, 9:42 AM UTC
Wand Ai logo

Staff Site Reliability Engineer

Wand Ai

🌍Probably WorldwideMay 28, 2026, 4:11 AM UTC
Madiffpl logo

DevOps Engineer – Ensure Scalability and Reliability

Madiffpl

🌍Probably WorldwideMay 28, 2026, 3:11 AM UTC
PlayOn logo

Senior Site Reliability Engineer

PlayOn

🌍Probably WorldwideMay 28, 2026, 12:29 AM UTC
Latitudesh Jobs logo

Senior Site Reliability Engineer

Latitudesh Jobs

🌍Probably WorldwideMay 28, 2026, 12:14 AM UTC
Icanbwell logo

Senior Site Reliability Engineer

Icanbwell

🌍Probably WorldwideMay 27, 2026, 10:08 PM UTC