LG

Senior Site Reliability Engineer (Linux, Python, Prometeus)

Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 27, 2026
Is this job info correct?


About the Role

We are seeking a Senior Site Reliability Engineer (SRE) to drive reliability, performance, and operational excellence across our enterprise virtualization platforms and high-performance AI compute infrastructure. In this role, you will act as a technical leader and mentor, automating away operational toil, elevating system observability, and designing resilient, self-healing infrastructure at scale.

This position is ideal for an experienced Linux Systems Specialist / DevOps Engineer who thrives on architecting robust systems, optimizing containerized workloads, and guiding engineering teams toward modern SRE best practices.

Key Responsibilities

  • Infrastructure Automation & Toil Reduction: Design, build, and maintain automation tooling and scripts (using Python, Golang, and IaC) to streamline software deployments, safeguard release processes, and minimize manual intervention.
  • Observability & Performance Optimization: Define Service Level Objectives (SLOs) and enhance system telemetry using Prometheus, Grafana, and distributed tracing to accelerate error detection and improve the reliability of our core virtualization platform.
  • AI Infrastructure & Capacity Management: Drive capacity planning, workload scheduling, and auto-scaling strategies tailored for high-demand AI compute infrastructure.
  • Incident Management & Reliability: Participate in on-call rotations, serving as a technical guide during critical service-impacting incidents to speed up recovery and root-cause resolution.
  • Technical Leadership & Mentorship: Elevate the engineering team by coaching developers on SRE principles, promoting a culture of reliability, and fostering continuous learning.

Required Qualifications & Technical Expertise

  • Core Experience: Deep, expert-level background in Linux/Unix Systems Administration, SRE, or DevOps supporting large-scale, distributed production environments.
  • Containerization & Orchestration: Proven expertise in Kubernetes and orchestrating large-scale containerized systems.
  • Software Engineering & IaC: Advanced proficiency in at least one modern programming language (Python or Golang) alongside configuration and IaC tools (Terraform, SaltStack, or Ansible).
  • Observability & SLOs: Hands-on experience establishing SLOs/SLIs and implementing enterprise observability stacks (Prometheus, Grafana, tracing frameworks).
  • System Architecture & Design: Practical experience architecting software systems and cloud/bare-metal infrastructure at scale.
  • SRE Evangelism: Strong accountability for system health, with a collaborative mindset to introduce SRE practices to teams new to the discipline.


Similar jobs

Apply for this job