Job Description This is a remote position. We are looking for a Site Reliability Engineer to join our team and help drive the reliability, scalability, security, and automation of modern cloud platforms. In this role, you will be responsible for operating and improving production environments, automating infrastructure management, and ensuring platform resilience across cloud and hybrid environments. You will work closely with Engineering, Security, and Platform teams to reduce operational overhead, improve reliability, and accelerate cloud transformation initiatives. The role combines cloud engineering, infrastructure automation, security, and operational excellence, with a strong focus on Microsoft Azure and Infrastructure as Code practices. Responsibilities: Collaborate with internal and external stakeholders to ensure the successful delivery of infrastructure and platform initiatives Deploy, maintain, and scale cloud and hybrid infrastructure environments Build, secure, and operate scalable cloud platforms, primarily in Microsoft Azure Manage compute, networking, and storage resources across production environments Partner with security teams to: Identify vulnerabilities Implement remediation actions Deploy security controls and endpoint protection solutions Ensure compliance with security standards Administer and optimize identity and access management solutions, including: Microsoft Entra ID (Azure AD) SSO configurations User permissions and access controls Develop and maintain Infrastructure as Code (IaC) solutions using Terraform Automate provisioning, configuration management, and operational recovery processes Implement and maintain monitoring and logging solutions to ensure high availability and rapid incident resolution Support FinOps initiatives through cost awareness, resource tagging, and governance practices Participate in an on-call rotation to ensure platform reliability and operational continuity Requirements Proven experience as a: Site Reliability Engineer (SRE) DevOps Engineer Systems Engineer supporting cloud production environments Strong operational experience with public cloud platforms, particularly: Microsoft Azure Experience with identity and access management technologies: Microsoft Entra ID (Azure AD) Single Sign-On (SSO) User and permission management Experience implementing and managing security tools such as: Endpoint Detection & Response (EDR) Vulnerability scanners Strong experience with: Terraform Infrastructure as Code (IaC) Infrastructure automation Experience with: Ansible Configuration management Server provisioning and patching Strong troubleshooting and problem-solving skills Ability to work autonomously while collaborating effectively with cross-functional teams Excellent written and verbal communication skills Fluency in English Nice-to-have Experience with container orchestration platforms: Kubernetes Azure Kubernetes Service (AKS) Experience managing CI/CD pipelines using: GitLab GitLab CI/CD Familiarity with Atlassian tools: Jira Jira Service Management (JSM) Confluence Experience with observability platforms: Prometheus Grafana Loki Azure certifications, such as: AZ-104: Azure Administrator Associate If this sounds like you, share your CV with us and let’s talk!
SITE RELIABILITY ENGINEER
Blackfinconsult
Site Reliability Engineer (APAC)
pod network
Site Reliability Engineer (AI)
Madiffpl
Staff Site Reliability Engineer
Wand Synthesis AI Inc
DevOps Engineer – Ensure Scalability and Reliability
Madiffpl
Senior Site Reliability Engineer
Latitudesh Jobs