American Express logo

Associate-Tech Operations Engineering

Hiring from
United Kingdom
Work type
Hybrid
Posted
Sep 25, 2026
Is this job info correct?

The Enterprise Technology Services organization partners with every part of the American Express business to power the company’s growth and innovation with trust and efficiency, and drive competitive differentiation with speed. We support the delivery and operations of technology, digital, and data capabilities, platforms, and services globally. Specifically, our team is responsible for the company’s technology engineering, architecture, and infrastructure, providing 24x7 support to ensure an uninterrupted, high-quality experience for customers and colleagues. We also provide product management for core enterprise platforms, and lead technology risk and information security, enterprise data governance and platforms, digital product and design, and enterprise AI platforms on behalf of the company.

As an Associate Technology Operations Engineer, you will help maintain the availability, stability and resilience of business-critical production applications and technology services. This role requires strong Technology Operations experience, combined with a proven knowledge of Incident Management and Application Support.

You will be expected to take ownership of high-severity production incidents, leading Major Incident bridges and providing clear direction throughout the incident lifecycle. This includes coordinating and assisting technical teams with troubleshooting and diagnostics, facilitating effective communication and collaboration across multiple support teams, managing escalations and stakeholder updates, identifying appropriate recovery actions, and driving incidents through to timely service restoration. You will maintain focus and momentum during critical situations, ensuring actions are clearly owned, tracked and progressed while minimising business and customer impact.

  • Provide operational support for business-critical production applications, platforms and infrastructure, ensuring high levels of availability, stability and resilience, while troubleshooting issues across application, infrastructure, network and cloud environments.

  • Lead and coordinate Major Incident bridges, taking ownership of high-severity incidents and managing them end-to-end through triage, impact assessment, prioritisation, escalation, investigation, recovery and closure, with a clear focus on rapid service restoration.

  • Coordinate and assist Application, Infrastructure, Network, Cloud, Engineering, SRE and third-party teams during complex incidents, maintaining structure and momentum by ensuring actions have clear owners, progress is tracked and appropriate technical and management escalations are made.

  • Assess customer and business impact and provide clear, timely incident communications to technical teams, business stakeholders and senior leadership, while maintaining accurate incident timelines, actions, decisions, impact assessments and restoration details

  • Facilitate and contribute to post-incident reviews and Root Cause Analysis (RCA), identifying corrective actions, reviewing incident trends and recurring issues, and driving opportunities to improve service stability, resilience and prevent recurrence.

  • Develop and maintain operational procedures, support documentation, runbooks and escalation processes, collaborating with Product, Engineering and business teams to continually improve operational readiness, monitoring, resilience and production support.

  • Act as an escalation point for Technology Operations Engineers, providing guidance during complex or high-severity incidents, removing blockers, sharing operational best practices, and supporting the development of less experienced engineers to promote consistent and effective production support.

Minimum Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering or comparable professional experience.
  • 4+ years of relevant experience in Technology Operations, Production Support, Service Management, SRE or a comparable operational environment.
  • Experience working in a high-volume, 24x7 production financial services environment, supporting business-critical applications and services with demanding availability, stability, resilience and operational requirements.
  • Strong hands-on Incident and Major Incident Management experience.
  • Proven experience leading Major Incident bridges for high-severity production incidents.
  • Demonstrated ability to coordinate multiple technical teams and drive incidents through service restoration.
  • Knowledge of ITIL/IT Service Management standards, particularly Incident, Major Incident, Problem and Change Management.
  • Experience communicating technical incidents and business impact to both technical and non-technical stakeholders.
  • Knowledge of Software Delivery Lifecycle and Disaster Recovery.
  • Strong troubleshooting experience across applications, infrastructure, networks, and/or cloud services.
  • Strong verbal and written communication, analytical and problem-solving skills.

Preferred Qualifications

  • Experience within financial services or another large-scale technology environment.
  • Experience with enterprise incident management, monitoring and observability tooling.
  • Experience with AWS, Azure, and/or Google Cloud.
  • Understanding of SRE practices and production resilience.
  • Knowledge of Python, Bash or automation technologies.
  • Relevant IT Service Management, cloud or technology operations certifications are advantageous.

Non-considerations for sponsorship:

Employment eligibility to work with American Express in the UK is required as the company will not pursue visa sponsorship for these positions.

Similar jobs

Apply for this job