Iscale Solutions logo

Lead Site Reliability Engineer

Hiring from
Philippines
Work type
Remote
Posted
Is this job info correct?
Show job description

Job Description

This is a remote position.

  • Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts.

  • Develop systems that are resilient to failures and ensure 99.9%+ uptime for critical services.

  • Lead incident response and post-incident reviews (blameless postmortems), ensuring robust root cause analysis and continuous improvement of systems.

  • Automate incident detection and response using automated runbooks or predefined workflows.

  • Write software as needed to support reliability or efficiency needs.

  • Design and implement full observability across systems using modern tools like Open Telemetry for tracing, metrics, and logging.

  • Use capacity planning, forecasting, and performance testing to ensure that the systems scale effectively as the user base and load grow.

  • Collaborate with development and operations teams on building reliable, scalable, and high-performance services.

  • Ensure best practices are followed across infrastructure design, deployment, and maintenance using tools like AWS, Kubernetes, EKS, Fargate, etc.

  • Champion Infrastructure as Code (IaC) to provision, manage, and scale infrastructure using tools like Terraform, Pulumi, or similar.

  • Get involved in chaos engineering initiatives.

  • Participate on our on-call rotation.

  • Drive advanced alerting and anomaly detection applied to metrics



Requirements

  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments.

  • Deep understanding of Linux systems, networking, and systems administration.

  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools.

  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki.

  • Strong skills in at least one programming language (Python, Go) to write production level code.

  • Strong skills in shell scripting using bash or similar. • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration.

  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc.



Benefits

  • Competitive Salary Package: Receive a pay package that matches your skills and experience.

  • Vacation and Sick Leave credits: Enjoy vacation and sick leave credits to maintain work-life balance.

  • Health Coverage: Get medical, dental, and vision insurance for you and your dependents.

  • Government-Mandated Benefits: Full coverage of all statutory benefits like SSS, PhilHealth, and Pag-IBIG.

  • Learning Opportunities: Access training, certifications, and mentorship to grow your career.

  • Team Engagement: Join team-building activities and wellness programs.

  • Modern Tools: Use the latest technology to excel in your role.

  • Career Growth: Clear paths for promotion and professional development.

  • Inclusive Culture: Be part of a diverse, supportive, and collaborative global team.

  • Referral Rewards: Earn bonuses for bringing great talent to the team.



Similar jobs

Apply for this job