Site Reliability Engineer (SRE)
- Hiring from
- United States
- Work type
- Remote
- Posted
Is this job info correct?
522,528 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Duration: 12+ Months
Location: 100% Remote
We are seeking a Senior Site Reliability Engineer (SRE) with 10+ years of experience to own the reliability, availability, observability, and production operations of a complex, multi-region global e-commerce platform. The environment is primarily AWS-based, with extensive use of ECS/Docker, Lambda, API Gateway, DynamoDB, CloudFront/Akamai, and SaaS platforms such as VTEX. This is a highly hands-on role that requires deep expertise in production incident management, observability, performance engineering, AWS troubleshooting, and resilience.
Key Responsibilities:
- Lead P1/P2 production incidents, drive recovery, and conduct RCA and blameless postmortems.
- Own observability using New Relic, AWS CloudWatch, and Amazon Athena.
- Define and monitor SLIs, SLOs, SLAs, reliability KPIs, and alerting strategies.
- Perform capacity planning, load testing, stress testing, and performance analysis using tools such as k6 and JMeter.
- Validate backup/restore, disaster recovery, resilience, patching, runtime upgrades, and SSL/TLS certificates.
- Troubleshoot AWS networking, ALB traffic, CloudFront/CDN, DNS, WAF, DDoS/Bot attacks, and Palo Alto firewall interactions.
- Troubleshoot production issues involving ECS, Lambda, S3, DynamoDB, API Gateway, and VTEX/SaaS integrations.
- Read and assess Terraform/IaC, Docker containers, and production infrastructure for reliability risks.
- Maintain operational runbooks and technical documentation in Confluence.
- 10+ years in SRE, Production Engineering, DevOps, or similar roles.
- Strong hands-on experience with AWS: VPC, ALB, CloudFront, ECS, S3, Lambda, WAF, Secrets Manager.
- Expert-level New Relic, CloudWatch, and Athena experience.
- Strong incident management, RCA, postmortems, SLI/SLO/SLA, and reliability engineering experience.
- Hands-on k6/JMeter performance and load testing.
- Strong knowledge of Terraform, Docker, Git, Linux/Unix, DNS, TLS/SSL, and shell scripting.
- Experience troubleshooting Node.js, PHP, Python, and Bash-based applications.
- Experience with CDN/edge technologies, traffic analysis, security, and production troubleshooting.
- Excellent communication and ability to operate independently in a mission-critical environment.
#dice