Site Reliability Engineer
- Hiring from
- Mexico
- Work type
- Remote
- Posted
Is this job info correct?
Show job descriptionHide job description
We are looking for a Site Reliability Engineer to keep cloud services reliable, observable, and automated across multi-tenant Kubernetes environments on Azure. You will own production health, reduce toil through automation, and partner with developers to improve resilience.
Responsibilities
- Operate Kubernetes clusters and containerized workloads running on Azure
- Troubleshoot production incidents end-to-end across network, OS, platform, and application layers
- Automate repetitive operational tasks with Python, Bash, or PowerShell to eliminate toil
- Define and track SLIs/SLOs and drive improvements to meet reliability targets
- Build and tune monitoring and alerting to detect issues before clients are impacted
- Improve platform reliability through capacity, performance, and failure-mode analysis
- Partner with development teams to harden services and improve operability standards
- Document runbooks and operational procedures to speed up diagnosis and recovery
- Perform root cause analysis and implement preventive actions after incidents
Requirements
- 2+ years of experience in Site Reliability Engineering or DevOps for production systems
- 2+ years of experience operating Kubernetes and containerized workloads
- Hands-on experience with Microsoft Azure services for running workloads
- SLA/SLO adherence experience including defining and tracking SLIs/SLOs
- Infrastructure fundamentals in networking and operating systems
- Strong Linux administration skills
- Strong scripting skills in Python, Bash, or PowerShell
- Incident response skills with calm, structured troubleshooting under pressure
- Clear communication skills for cross-team collaboration during incidents and reviews
- Collaborative mindset to work effectively with development teams
- English proficiency level B2 (Upper-Intermediate)
Nice to have
- Argo CD experience for GitOps-based deployments
- Elastic Stack experience for observability workflows
- Istio experience for service mesh traffic management
- Windows Administration experience including Windows Server operations
We offer
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.
Remote in Mexico