The Brixton Group logo

Site Reliability Engineer (SRE)

Hiring from
United States
Work type
Remote
Posted
Is this job info correct?

522,528 remote jobs, straight from company career pages

100% free · New jobs every hour

Show job description

Duration: 12+ Months
Location: 100% Remote

We are seeking a Senior Site Reliability Engineer (SRE) with 10+ years of experience to own the reliability, availability, observability, and production operations of a complex, multi-region global e-commerce platform. The environment is primarily AWS-based, with extensive use of ECS/Docker, Lambda, API Gateway, DynamoDB, CloudFront/Akamai, and SaaS platforms such as VTEX. This is a highly hands-on role that requires deep expertise in production incident management, observability, performance engineering, AWS troubleshooting, and resilience.

Key Responsibilities:
  • Lead P1/P2 production incidents, drive recovery, and conduct RCA and blameless postmortems.
  • Own observability using New Relic, AWS CloudWatch, and Amazon Athena.
  • Define and monitor SLIs, SLOs, SLAs, reliability KPIs, and alerting strategies.
  • Perform capacity planning, load testing, stress testing, and performance analysis using tools such as k6 and JMeter.
  • Validate backup/restore, disaster recovery, resilience, patching, runtime upgrades, and SSL/TLS certificates.
  • Troubleshoot AWS networking, ALB traffic, CloudFront/CDN, DNS, WAF, DDoS/Bot attacks, and Palo Alto firewall interactions.
  • Troubleshoot production issues involving ECS, Lambda, S3, DynamoDB, API Gateway, and VTEX/SaaS integrations.
  • Read and assess Terraform/IaC, Docker containers, and production infrastructure for reliability risks.
  • Maintain operational runbooks and technical documentation in Confluence.
Required Skills:
  • 10+ years in SRE, Production Engineering, DevOps, or similar roles.
  • Strong hands-on experience with AWS: VPC, ALB, CloudFront, ECS, S3, Lambda, WAF, Secrets Manager.
  • Expert-level New Relic, CloudWatch, and Athena experience.
  • Strong incident management, RCA, postmortems, SLI/SLO/SLA, and reliability engineering experience.
  • Hands-on k6/JMeter performance and load testing.
  • Strong knowledge of Terraform, Docker, Git, Linux/Unix, DNS, TLS/SSL, and shell scripting.
  • Experience troubleshooting Node.js, PHP, Python, and Bash-based applications.
  • Experience with CDN/edge technologies, traffic analysis, security, and production troubleshooting.
  • Excellent communication and ability to operate independently in a mission-critical environment.
26-01122
#dice

Similar jobs

Apply on LinkedIn