AC
Site Reliability Engineer (SRE)(NO OPT/CPT) / Remote
- Hiring from
- United States
- Work type
- Remote
- Posted
Is this job info correct?
Show job descriptionHide job description
Position: Site Reliability Engineer (SRE)(NO OPT/CPT)
Location: NY/NJ- Remote
Employment Type: Full Time
Job Description-
We are looking for an experienced Site Reliability Engineer (SRE) to help build, maintain, and improve highly available, scalable, and reliable production systems. The ideal candidate will have strong experience with cloud infrastructure, automation, monitoring, incident management, and DevOps practices.
You will work closely with software engineering, DevOps, security, and infrastructure teams to improve system reliability, automate operational processes, and ensure excellent application performance and availability.
Key Responsibilities
Location: NY/NJ- Remote
Employment Type: Full Time
Job Description-
We are looking for an experienced Site Reliability Engineer (SRE) to help build, maintain, and improve highly available, scalable, and reliable production systems. The ideal candidate will have strong experience with cloud infrastructure, automation, monitoring, incident management, and DevOps practices.
You will work closely with software engineering, DevOps, security, and infrastructure teams to improve system reliability, automate operational processes, and ensure excellent application performance and availability.
Key Responsibilities
- Design, implement, and maintain highly available and scalable production environments.
- Develop and maintain automation for infrastructure provisioning, deployment, monitoring, and operational tasks.
- Manage and optimize cloud infrastructure across AWS, Azure, or GCP.
- Build and maintain CI/CD pipelines to support reliable and frequent software releases.
- Implement Infrastructure as Code using tools such as Terraform, CloudFormation, or Ansible.
- Monitor system health, availability, performance, and capacity using tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
- Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews.
- Establish and maintain SLIs, SLOs, and SLAs for critical applications and services.
- Identify recurring operational issues and develop automation to eliminate manual processes.
- Collaborate with development teams to improve application reliability, scalability, and performance.
- Support containerized and microservices-based environments using Docker and Kubernetes.
- Implement and maintain logging, alerting, observability, and performance-monitoring solutions.
- Contribute to disaster recovery, business continuity, and capacity-planning initiatives.
- Maintain documentation for infrastructure, operational procedures, architecture, and incident-response processes.
- Participate in an on-call rotation and respond to production incidents when required.
- 4+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Infrastructure Engineering.
- Strong experience with Linux/Unix systems administration.
- Hands-on experience with at least one major cloud platform: AWS, Azure, or GCP.
- Strong scripting/programming skills in Python, Bash, Go, or similar languages.
- Experience with Kubernetes and Docker.
- Strong knowledge of CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Experience with Terraform or other Infrastructure-as-Code technologies.
- Hands-on experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
- Strong understanding of networking, DNS, HTTP/HTTPS, load balancing, and security fundamentals.
- Experience troubleshooting complex production issues and performing root-cause analysis.
- Strong understanding of SRE principles, reliability engineering, automation, and incident management.
- Excellent communication and problem-solving skills.
- Experience supporting large-scale distributed systems and microservices.
- Experience with service mesh technologies such as Istio or Linkerd.
- Knowledge of Helm, ArgoCD, Flux, or other Kubernetes deployment tools.
- Experience with cloud-native architectures and serverless technologies.
- Familiarity with security and compliance practices in cloud environments.
- Experience implementing observability using OpenTelemetry.
- Experience with performance testing, capacity planning, and chaos engineering.
- Relevant cloud certifications such as AWS, Azure, or Google Cloud certifications.