Sr Engineer, Site Reliability Engineer
- Salary
- $111.8K–$134.2K
- Hiring from
- United States
- Work type
- Hybrid
- Posted
514,029 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Job Purpose
In this role, you will serve as a senior technical leader responsible for designing, building, and evolving reliability engineering practices across critical platforms and services. You will partner closely with Engineering, Architecture, Security, and Product teams to improve system reliability, scalability, resiliency, and performance through automation, observability, and engineering excellence.
You will establish reliability standards, define SLOs and SLIs, drive adoption of reliability best practices, and leverage data-driven insights to continuously improve customer experience and platform health. As a technical expert, you will influence architecture decisions, lead reliability initiatives, and help engineering teams build highly available, fault-tolerant systems at scale.
Location: St. Louis, MO (hybrid)
Responsibilities
- Design, implement, and evolve Site Reliability Engineering practices for mission-critical applications and distributed systems.
- Define, implement, and govern Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to measure and improve service reliability.
- Drive reliability improvements through automation, self-healing capabilities, resiliency engineering, and reduction of operational toil.
- Partner with engineering teams to embed reliability, scalability, security, and observability requirements into system architecture and software development lifecycles.
- Develop and maintain Infrastructure as Code (IaC), platform automation, CI/CD pipelines, and reliability tooling to improve deployment consistency and operational efficiency.
- Lead reliability reviews, architecture assessments, and resiliency testing efforts, including failure mode analysis, fault injection, and chaos engineering practices.
- Design and implement comprehensive observability solutions utilizing metrics, logs, traces, synthetic monitoring, and real user monitoring.
- Analyze production performance, reliability trends, and failure patterns to proactively identify systemic risks and drive long-term improvements.
- Collaborate with development teams to optimize application performance, scalability, resource utilization, and cloud infrastructure efficiency.
- Provide technical leadership during complex production incidents by supporting root cause analysis and identifying opportunities for reliability improvements.
- Drive adoption of engineering best practices for high availability, disaster recovery, fault tolerance, and business continuity.
- Establish reliability standards, reference architectures, engineering patterns, and platform capabilities to enable scalable and resilient service delivery.
- Mentor engineers and act as a subject matter expert in SRE, cloud architecture, observability, automation, and distributed systems engineering.
Requirements
- 10+ years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Infrastructure Engineering, or DevOps, supporting large-scale, highly available systems.
- Deep experience designing, operating, and optimizing cloud-native and distributed systems in public cloud environments (Preferably GCP).
- Strong expertise in reliability engineering concepts, including SLOs, SLIs, error budgets, availability modeling, scalability, resiliency, and performance engineering.
- Hands-on experience building infrastructure and platform automation using Infrastructure as Code tools such as Terraform or CloudFormation.
- Strong experience with Kubernetes, container platforms, orchestration technologies, and modern cloud-native architectures.
- Experience implementing observability strategies utilizing monitoring, logging, tracing, and application performance management platforms.
- Strong software engineering and scripting skills using languages such as Python, Go, Java, Bash, or similar.
- Experience building and supporting CI/CD pipelines and modern software delivery practices.
- Strong understanding of networking, operating systems, security principles, databases, and application architectures.
- Demonstrated ability to perform deep technical troubleshooting and root cause analysis in complex distributed environments.
- Excellent communication and collaboration skills, with proven ability to influence technical decisions across multiple engineering teams.
Preferred Qualifications
- Experience designing and operating large-scale distributed systems supporting millions of transactions or users.
- Strong background in observability and APM platforms such as Dynatrace, Datadog, New Relic, Splunk, Grafana, OpenTelemetry, AppDynamics, or equivalent technologies.
- Experience implementing advanced reliability practices such as chaos engineering, fault injection testing, resiliency validation, and performance benchmarking.
- Deep expertise in cloud platform architecture, including compute, networking, storage, service mesh, and container ecosystem technologies.
- Experience building internal developer platforms, self-service infrastructure capabilities, and platform engineering solutions.
- Knowledge of disaster recovery, business continuity, and multi-region architecture design, including RTO/RPO strategies.
- Experience operating in regulated environments with security and compliance requirements such as PCI, SOX, SOC2, or ISO 27001.
- Demonstrated technical leadership and ability to influence engineering strategy, architecture decisions, and organizational adoption of SRE principles.
- Experience leveraging OpenTelemetry and modern observability frameworks to correlate business, application, and infrastructure telemetry.
- Passion for automation, engineering excellence, continuous improvement, and building resilient systems at enterprise scale.
Additional Description :