Lead Site Reliability Engineer
- Hiring from
- Philippines
- Work type
- Remote
- Posted
Show job descriptionHide job description
Job Description
This is a remote position.
-
Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.
-
Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.
-
Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.
-
Automate incident detection and response using automated runbooks or predefined workflows.
-
Write software as needed to support reliability or efficiency needs.
-
Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.
-
Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.
-
Collaborate with development and operations teams on building reliable, scalable, and high-performance services.
-
Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.
-
Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.
-
Get involved in chaos engineering initiatives.
-
Participate on our on-call rotation.
-
Drive advanced alerting and anomaly detection applied to metrics
Requirements
-
3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.
-
Deep understanding of Linux systems, networking, and systems administration.
-
Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.
-
Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.
-
Strong skills in at least one programming language (Python, Go) to write production level code.
-
Strong skills in shell scripting using bash or similar. • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.
-
Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.
Benefits
-
Competitive Salary Package: Receive a pay package that matches your skills and experience.
-
Vacation and Sick Leave credits: Enjoy vacation and sick leave credits to maintain work-life balance.
-
Health Coverage: Get medical, dental, and vision insurance for you and your dependents.
-
Government-Mandated Benefits: Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.
-
Learning Opportunities: Access training, certifications, and mentorship to grow your career.
-
Team Engagement: Join team-building activities and wellness programs.
-
Modern Tools: Use the latest technology to excel in your role.
-
Career Growth: Clear paths for promotion and professional development.
-
Inclusive Culture: Be part of a diverse, supportive, and collaborative global team.
-
Referral Rewards: Earn bonuses for bringing great talent to the team.