As a Staff Engineer – Site Reliability Engineering, you will play a key technical leadership role in designing, building, and evolving our hybrid infrastructure and developer platforms. You’ll work hands-on across cloud automation, high-performance bare-metal systems, and DevOps tooling to deliver reliable, scalable infrastructure that accelerates software delivery. This role is highly cross-functional. You will collaborate closely with Development, DevOps, and QA teams to solve infrastructure and release challenges, support delivery cycles, and continuously improve platform capabilities. You’ll balance strong technical execution with architectural influence, contributing to system design while remaining directly involved in implementation and operational support. Occasional in-person meetings or team events may be required. Key Responsibilities Infrastructure & Platform Engineering Ensure system availability and reliability through automated monitoring strategies Produce post-mortems and implement resulting process improvements Proactively mitigate operational risks through risk assessment and wider collaboration with engineering teams Design, implement, and continuously measure and improve risk mitigation strategies Monitor system health through observability and telemetry Unblock bottlenecks in system performance Minimize emergency response time priods Maintain internal tooling surrounding bug tracking, CI/CD pipelines, and wider communication with the teams Engineering Collaboration & Delivery Support Partner with Dev, DevOps, and QA teams to resolve infrastructure or deployment blockers during release cycles Provide technical guidance and mentorship to platform engineers Participate in architectural reviews, release readiness checkpoints, and root-cause analyses Required Qualifications 10+ years of experience in infrastructure, platform, or DevOps engineering roles Strong programming skills (e.g., Python, Go) Hands-on experience with hybrid infrastructure (cloud + bare metal) Deep knowledge of infrastructure-as-code tools (Terraform, Helm, Kubernetes) Proven cross-functional collaboration skillsPreferred Qualifications Experience with GPU-accelerated compute or HPC-style infrastructure Familiarity with platform engineering or developer experience optimization Experience with high-velocity release cycles and incident response Preferred Qualifications Experience with GPU-accelerated compute or HPC-style infrastructure Familiarity with platform engineering or developer experience optimization Experience with high-velocity release cycles and incident response
Senior Site Reliability Engineer (SRE) I
Thomsonreuters
Systems Development Engineer (SRE/DevOps)
Model N
Site Reliability Engineer (SRE)
TTEC Digital
Sr Manager Engineering - SRE
Reltio
Engineer III - TechOps CICD SRE (Reliability Focused)
Crowdstrike
Staff Software Engineer I - SRE
Confluent