Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Yum! logo

Site Reliability Engineer III

Yum!
Posted 3 weeks ago
🇻🇳Vietnam🏢Hybrid📁Engineering & Development
Is this job info correct?

The Site Reliability Engineer (Level 7) is an experienced mid-level individual contributor responsible for the reliability, scalability, performance, and operational excellence of one or more markets, platforms, or critical services. Incident Management and Reliability Independently lead complex incidents involving multiple systems, teams, or dependencies. Coordinate incident response activities, facilitate communication, and drive timely resolution. Lead or contribute to post-incident reviews and root cause analysis activities. Ensure corrective and preventive actions are identified, prioritized, tracked, and completed Monitoring, Alerting, and Observability Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions. Develop meaningful alerts based on service behavior, customer impact, and business priorities. Build and maintain dashboards that provide actionable insights into system performance and reliability. Platform and Market Ownership Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end. Establish and maintain operational excellence standards for assigned domains. Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date. Continuously assess platform health, identify reliability risks, and drive improvements before incidents occur. Partner with engineering teams to ensure new features and services meet reliability requirements before production release. Define and track reliability metrics, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets. Automation, and AI Develop and maintain tools, scripts, and automation solutions that reduce manual effort and improve operational efficiency. Identify and eliminate repetitive tasks through automation and self-service capabilities. Establish and promote best practices for the responsible use of AI within SRE workflows Team Contribution and Mentoring Mentor and support Level 5 and Level 6 engineers in incident management, monitoring, automation, AI adoption, and operational best practices. Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards. Share knowledge through training sessions, documentation, and post-incident learning activities. Contribute to the continuous improvement of team processes, standards, and ways of working. Own SRE responsibilities for one or more markets, platforms, or critical services end-to-end. Establish and maintain operational excellence standards for assigned domains. Ensure monitoring coverage, dashboards, runbooks, and alerting configurations remain accurate, effective, and up to date. Coordinate incident response activities, facilitate communication, and drive timely resolution. Design, implement, and continuously optimize monitoring, logging, alerting, and tracing solutions. Develop meaningful alerts based on service behaviours, customer impact, and business priorities. Review monitoring configurations, dashboards, runbooks, and operational documentation to maintain quality standards. Lead or contribute to post-incident reviews and root cause analysis activities. Influence technical decisions that improve platform stability, scalability, and operational efficiency

Similar jobs

Similar jobs

Sleek logo

Senior Site Reliability Engineer (SRE)

Sleek

🌍India, Vietnam4 weeks ago
GRADION logo

Site Reliability Engineer (SRE)/ DevOps Engineer (Middle - Senior)

GRADION

🌍Thailand, VietnamJul 2, 2026, 7:12 AM UTC
Global Fashion Group SGP Services PTE Ltd. logo

Senior Site Reliability Engineer

Global Fashion Group SGP Services PTE Ltd.

🇻🇳VietnamJul 2, 2026, 7:12 AM UTC
GRADION logo

Senior DevOps/ Site Reliability Engineer (AWS/GCP)

GRADION

🌍Thailand, VietnamJun 18, 2026, 7:53 AM UTC
Global Fashion Group SGP Services PTE Ltd. logo

Site Reliability Engineer

Global Fashion Group SGP Services PTE Ltd.

🇻🇳VietnamMay 28, 2026, 12:31 AM UTC
QV

Senior/Principal JavaScript Engineer

Qode.World Vietnam

🇻🇳Vietnam9 hours ago