LG

Senior SRE (ACDC Platform / Infrastructure Automation)

Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 27, 2026
Is this job info correct?

What You Will Do:

  • Architect and Build Automation Solutions: Your primary focus will be on designing and developing robust automation using Ansible. You will create playbooks, roles, and workflows that manage a massive, enterprise-scale infrastructure, effectively eliminating operational toil.
  • Own the Infrastructure as Code (IaC) Lifecycle: You will manage and enhance IaC solutions using tools like Terraform, Ansible, and SaltStack, ensuring that the entire infrastructure is versioned, repeatable, and scalable.
  • Engineer for Ultimate Reliability: You will design, develop, and operate essential backend services, with a relentless focus on improving reliability, scalability, and performance. This includes building proactive observability solutions (monitoring, logging, alerting) to detect and resolve issues before they impact customers.
  • Become a Technical Leader and Mentor: You will act as a subject matter expert for the core platform systems, guiding other engineers and sharing best practices in automation and reliability across the organization.
  • Solve the Toughest Problems: You will collaborate with software, infrastructure, and platform teams to troubleshoot complex production issues, acting as a key technical resource during incident response and participating in an on-call rotation.

Who We're Looking For:

  • You are an expert in Ansible. You have demonstrable, hands-on experience developing playbooks, creating custom roles, and managing complex, enterprise-scale configurations.
  • You have a deep understanding and practical experience with Infrastructure as Code (IaC) principles and tools such as Terraform, SaltStack, Chef, or Puppet.
  • You possess advanced, expert-level skills in Linux engineering and administration, with a proven ability to design, build, and deploy software and infrastructure at a massive scale.
  • You have a strong background as a Site Reliability Engineer or Software Engineer, with significant experience working on large-scale, distributed systems.
  • You are not just a user of tools, but a builder. You have a track record of developing automation and internal tools to improve reliability and operational efficiency.
  • You are an excellent communicator and collaborator, with the ability to work effectively across teams and mentor others on SRE best practices.


Similar jobs

Apply for this job