Site Reliability Engineer I
- Hiring from
- India
- Work type
- Hybrid
- Posted
Is this job info correct?
521,292 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.
- Collaborates with Software Engineering teams to support the development, and implementation of features that enhance system resilience, scalability, and performance, ensuring systems can handle varying loads and recover effectively from unexpected disruptions
- Collaborates in the development and implementation of automation tools and frameworks, including infrastructure as code (IaC) practices, to reduce manual intervention and improve system efficiency, with guidance from peers and leaders
- Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
- Collaborates in the design and execution of chaos engineering experiments and other resiliency testing methods to help identify potential failure points and ensure systems can recover from disruptions, with guidance from peers and leaders
- Supports the development and implementation of disaster recovery plans and business continuity strategies, ensuring systems can recover quickly and effectively from unexpected disruptions
- Collaborates with seniors to promote and implement best practices such as error budgeting, service-level objectives (SLOs), and service-level indicators (SLIs), contributing to a culture of continuous improvement and reliability
- Collaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives
Education Qualifications:
- Bachelor’s degree in Computer Science, Information Technology, Engineering, and/or comparable experience; advance degree preferred
- Knowledge of modern observability stack – Splunk, Elastic Search, Prometheus, Grafana
- Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture
- Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms
- Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud
Work Experience:
- Experience in software development, or technology operations, with a focus on Site Reliability Engineering
- Experience in Linux/Unix systems, object-oriented programming languages (e.g., Java), scripting languages (e.g., Python, Bash), and cloud platforms (e.g., AWS, Azure, GCP)
Licenses and Certifications:
- Advanced certification in Site Reliability Engineering (SRE) or related is a plus