The Enterprise Technology Services organization partners with every part of the American Express business to power the company’s growth and innovation with trust and efficiency, and drive competitive differentiation with speed. We support the delivery and operations of technology, digital, and data capabilities, platforms, and services globally. Specifically, our team is responsible for the company’s technology engineering, architecture, and infrastructure, providing 24x7 support to ensure an uninterrupted, high-quality experience for customers and colleagues. We also provide product management for core enterprise platforms, and lead technology risk and information security, enterprise data governance and platforms, digital product and design, and enterprise AI platforms on behalf of the company. Manager, Site Reliability Engineering leads and mentors Site Reliability Engineering (SRE) teams, fostering a culture of continuous improvement and inclusivity, while collaborating across the organization to enhance system resilience, scalability, and alignment with business objectives. Manages and leads a team of Site Reliability Engineering colleagues, enabling a culture of continuous learning, growth opportunities, and inclusivity for all individual colleagues and teams Provides leadership, guidance, and coaching to Site Reliability Engineering teams, supporting training and development of best practices in software development, resiliency, and non-functional system requirements Recruit and develop a high-performing team, recognizing and rewarding achievements, and creating an environment that motivates and energizes colleagues to achieve best business objectives Oversees and facilitates collaboration with Software Engineering teams to design and implement features that improve system resilience, scalability, and performance; ensuring optimal functionality Collaborates with executives, product managers, and other stakeholders to ensure SRE principles are embedded throughout the organization Leads comprehensive chaos engineering experiments and resiliency tests, driving the analyzation of outcomes and implementation of improvements that enhance system robustness and recovery capabilities Plans regular drills and strategic planning to ensure organization is prepared for and can swiftly recover from complex and unexpected disruptions Collaborates and co-creates effectively with teams in product and the business to align technology initiatives with business objectives. Balance feature development speed and reliability with well-defined service level objectives Recognizes opportunities to adopt innovative technologies to enable business capabilities. Explores new automation techniques to refine the agility, speed and quality of engineering initiatives and efforts Bachelor’s degree in computer science, Information Technology, Engineering, and/or comparable experience; advance degree preferred Knowledge of modern observability stack – Splunk, Elastic Search, Prometheus, Grafana Knowledge of containerization technologies (e.g., Kubernetes, Docker) and microservices architecture Knowledge of observability tools and methodologies, including experience with logging, monitoring, tracing, and performance analysis platforms Knowledge of cloud-based Site Reliability Engineering (SRE) practices and experience with public cloud platforms such as AWS, Azure, or Google Cloud. Knowledge of Jira, confluence, rally and project management tools including MS office suit. Work Experience: Must have Domain knowledge of Cards Payments systems. Understanding of E2E workflows of Authorization Approval and clearing & reconciliation processes. Have SRE experience with knowledge of SRE functions Have knowledge on Splunk, ELF/Kibana and Prometheus/Grafana and experience to use these tools to automate and configure alerting and build dashboards. Have experience in application support (Must have). Application support in cloud-based environment. Conceptual/support knowledge of microservices in cloud environment and deployment process Incident management system knowledge (service now) Good communication skills to run production bridges. Experience in source control using tools such as Git with DevOps and IT automation concepts. Basic UNIX knowledge and any programing language (preferably java, go lang or UI stack). Some knowledge in Redhat Open Shift 3.9/3.11 or Kubernetes 1.9/1.11 and above. Perform day today support activities to track incidents, respond timely on incidents and review and analyze issues at level 2. Create automation dashboards using Splunk, Grafana and Kibana Flexible to work shifts (only day shift, start may be little late than usual time). And ready to provide weekend support as per roster. Review current issues and work with Engineering team to get code fixed and deployed. Familiar with Agile or other rapid application development methods Experience with design and coding across one or more platforms and languages as appropriate Experience with distributed (multi-tiered) systems, algorithms, and relational databases A proactive approach to spotting problems, areas for improvement, and performance bottlenecks. Good Knowledge of Networking & Services like TCP/UDP, HTTPs, rest APIs. Able to understand and use complex data structures and associated components Designs, codes, tests, maintains, and documents applications Takes part in reviews of own work and reviews of colleagues' work Defines test conditions based on the requirements and specifications provided Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions.
TechOps Engineering Lead
Sparta Commodities
MLOps Engineering Manager
Trainline
Senior Director, Engineering- X-Ops Platform
Sophos
Senior Project Manager
Costain Group
Senior Consultant - Strategy3
Ipsos
Regional Operations Manager
Averna