Senior Site Reliability Engineer (Linux, Python, Prometeus)
- Hiring from
- Probably Worldwide
- Work type
- Remote
- Posted
- Sep 27, 2026
About the Role
We are seeking a Senior Site Reliability Engineer (SRE) to drive reliability, performance, and operational excellence across our enterprise virtualization platforms and high-performance AI compute infrastructure. In this role, you will act as a technical leader and mentor, automating away operational toil, elevating system observability, and designing resilient, self-healing infrastructure at scale.
This position is ideal for an experienced Linux Systems Specialist / DevOps Engineer who thrives on architecting robust systems, optimizing containerized workloads, and guiding engineering teams toward modern SRE best practices.
Key Responsibilities
- Infrastructure Automation & Toil Reduction: Design, build, and maintain automation tooling and scripts (using Python, Golang, and IaC) to streamline software deployments, safeguard release processes, and minimize manual intervention.
- Observability & Performance Optimization: Define Service Level Objectives (SLOs) and enhance system telemetry using Prometheus, Grafana, and distributed tracing to accelerate error detection and improve the reliability of our core virtualization platform.
- AI Infrastructure & Capacity Management: Drive capacity planning, workload scheduling, and auto-scaling strategies tailored for high-demand AI compute infrastructure.
- Incident Management & Reliability: Participate in on-call rotations, serving as a technical guide during critical service-impacting incidents to speed up recovery and root-cause resolution.
- Technical Leadership & Mentorship: Elevate the engineering team by coaching developers on SRE principles, promoting a culture of reliability, and fostering continuous learning.
Required Qualifications & Technical Expertise
- Core Experience: Deep, expert-level background in Linux/Unix Systems Administration, SRE, or DevOps supporting large-scale, distributed production environments.
- Containerization & Orchestration: Proven expertise in Kubernetes and orchestrating large-scale containerized systems.
- Software Engineering & IaC: Advanced proficiency in at least one modern programming language (Python or Golang) alongside configuration and IaC tools (Terraform, SaltStack, or Ansible).
- Observability & SLOs: Hands-on experience establishing SLOs/SLIs and implementing enterprise observability stacks (Prometheus, Grafana, tracing frameworks).
- System Architecture & Design: Practical experience architecting software systems and cloud/bare-metal infrastructure at scale.
- SRE Evangelism: Strong accountability for system health, with a collaborative mindset to introduce SRE practices to teams new to the discipline.