We are looking for a passionate and experienced DevOps Engineer to build, automate, and operate highly available cloud infrastructure that powers our business-critical applications and data platforms. The ideal candidate will play a key role in ensuring the 24x7 availability, reliability, security, and performance of production systems by leveraging AWS cloud technologies, automation, CI/CD, observability, and operational excellence. This role requires strong expertise in cloud infrastructure, deployment automation, production support, networking, monitoring, and incident management. You will work closely with Development, QA, Data Engineering, and Infrastructure teams to deliver scalable, resilient, and secure solutions while driving continuous improvement across the software delivery lifecycle. Cloud Infrastructure & Automation • Design, deploy, and manage secure, scalable, and highly available infrastructure on AWS. • Develop and maintain Infrastructure as Code (IaC) using AWS CloudFormation to automate infrastructure provisioning and configuration. • Provision, configure, and manage AWS services including Amazon ECS, Amazon RDS, Amazon DynamoDB (DDB), Application Load Balancers (ALB), Route 53, IAM, CloudWatch, S3, and AWS Secrets Manager. • Optimize cloud resources for performance, availability, scalability, and cost efficiency. • Implement infrastructure automation to improve operational consistency and reduce manual effort. CI/CD & Release Engineering • Design, implement, and maintain robust CI/CD pipelines using Bamboo and Bitbucket. • Automate application deployments, infrastructure changes, and release processes across development, testing, staging, and production environments. • Support release planning, deployment validation, rollback strategies, and change management activities. • Collaborate with development teams to improve deployment reliability and accelerate software delivery. Production Operations & Support • Ensure 24x7 operational availability of production applications, cloud infrastructure, and business-critical data pipelines. • Participate in on-call rotations, Production Governance (PG) calls, and major incident bridge calls. • Proactively monitor application health, infrastructure performance, and pipeline execution using Splunk, New Relic, and Amazon CloudWatch. • Respond to production incidents, outages, performance degradation, and infrastructure failures while meeting SLA commitments. • Lead incident troubleshooting, perform Root Cause Analysis (RCA), and implement corrective and preventive actions. • Execute emergency deployments, hotfixes, rollback procedures, and infrastructure recovery activities. • Develop and maintain operational runbooks, SOPs, and post-incident documentation. Data Pipeline Reliability • Monitor, maintain, and support business-critical ETL workflows and data pipelines. • Investigate pipeline failures and optimize pipeline performance and reliability. • Collaborate with Data Engineering teams to improve resilience, monitoring, and automation. • Ensure data platform availability and adherence to operational SLAs. Monitoring & Observability • Build monitoring, logging, and alerting solutions using Splunk, New Relic, and CloudWatch. • Create dashboards and alerts for infrastructure, application, and pipeline health. • Analyze metrics and logs to improve performance and operational efficiency. • Drive continuous improvements in observability and platform reliability. Networking & Cloud Architecture • Configure and manage AWS networking components including VPCs, Subnets, Route Tables, Security Groups, Network ACLs, NAT Gateways, Internet Gateways, and ALBs. • Manage and troubleshoot DNS using Amazon Route 53. • Troubleshoot routing, connectivity, SSL/TLS, DNS resolution, and load balancing issues. • Design secure, highly available, and fault-tolerant cloud architectures. High Availability & Disaster Recovery • Design and implement Multi-AZ and Multi-Region deployment strategies. • Support disaster recovery planning, failover testing, and resilience engineering. • Implement automated recovery mechanisms and continuously improve platform resiliency. Collaboration & Continuous Improvement • Partner with Development, QA, Infrastructure, Security, and Data Engineering teams. • Automate manual processes, improve operational efficiency, and reduce technical debt. • Promote DevOps culture and cloud best practices. • Stay current with emerging cloud technologies. Cloud & AWS • AWS CloudFormation • Amazon ECS • Amazon RDS • Amazon DynamoDB (DDB) • Application Load Balancer (ALB) • Amazon Route 53 • IAM • Amazon CloudWatch • Amazon S3 • AWS Secrets Manager DevOps & CI/CD • Bamboo • Bitbucket • Git • CI/CD pipeline development • Release Management • Deployment Automation • Infrastructure Automation Monitoring & Observability • Splunk • New Relic • CloudWatch • Log Analysis • Alerting & Dashboarding • Performance Monitoring Networking • TCP/IP • DNS • HTTP/HTTPS • SSL/TLS • AWS VPC • Route Tables • Security Groups • Load Balancing • NAT Gateway • VPN • Networking Troubleshooting Automation & Scripting • Bash • Python • Shell Scripting Operational Excellence • Production Support • Incident Management • Root Cause Analysis (RCA) • Change & Release Management • High Availability (HA) • Disaster Recovery (DR) • Multi-Region Architecture • Operational Monitoring Preferred Qualifications • Experience with Docker and containerized applications. • Familiarity with Kubernetes (Amazon EKS). • Experience with Terraform or other Infrastructure as Code tools. • Knowledge of AWS Well-Architected Framework. • Experience implementing blue-green, rolling, or canary deployment strategies. • AWS Certifications are highly desirable. Experience • 3–8+ years of experience in DevOps, Cloud Engineering, or Site Reliability Engineering. • Proven experience managing AWS production environments and enterprise-scale cloud infrastructure. • Strong experience supporting 24x7 production operations, mission-critical applications, and data platforms. • Experience participating in on-call rotations, Production Governance (PG) calls, incident management, and outage resolution. • Demonstrated ability to build reliable CI/CD pipelines, automate infrastructure, and improve platform observability and operational excellence. What You'll Bring • A strong ownership mindset with a passion for automation, reliability, and operational excellence.
Specialist, ESE DevOps Engineer
Msd
Senior Devops Engineer
Distro
Product Developer Senior - DevOps Engineer
Epicorsoftware
Sr. Engineer - DevOps 4B
Genpact
DevOps Engineer
Weekday
Production Support / DevOps Engineer - Azure Stack
Keyrus Group