This role is responsible for driving the reliability, resiliency, performance, and modernization of critical American Express platforms across Distributed environments. You will leverage deep technical expertise in software engineering, runtime engineering, production support, and platform operations to quickly assess and remediate complex availability, performance, and operational issues. As part of our technology team, you will partner with engineering, product, infrastructure, and operations teams to design, build, automate, and support highly available enterprise platforms. You will help accelerate modernization initiatives, improve operational excellence, and deliver secure, scalable, and resilient solutions that power critical customer and business capabilities. Software Engineering & Platform Development Serve as a hands-on engineer with experience in supporting complex enterprise applications, platforms, and operational tooling across Distributed environments. Design, develop, prototype, code, test, and implement scalable software solutions using technologies such as Java, Python, SQL, and related frameworks. Act as a technical contributor in change management reviews, root cause analysis, and troubleshooting of complex technical issues. Design and implement automation solutions, and engineering practices that improve platform resiliency, operational efficiency, and security. Use best practices in incident management, problem management and change management as this role focuses on application production support. Runtime Engineering, Reliability & Operations Contribute to the technical roadmap for runtime systems, ensuring platform reliability, scalability, availability, recoverability, and performance. Establish, monitor, and continuously improve key performance indicators (KPIs), service level objectives (SLOs), and operational metrics (MTTR, MTBF) for platform health and resiliency. Perform diagnosis and resolution of production incidents, batch failures, application outages, performance bottlenecks, and infrastructure issues across Mainframe and Distributed platforms. Apply Site Reliability Engineering (SRE) principles and operational excellence practices to improve system stability and reduce operational risk. Support disaster recovery, high availability, workload management, capacity planning, and business continuity initiatives. Distributed Platform Engineering Must have experience in Support and optimize enterprise platforms across distributed technologies including Java, Python, C, SQL, NodeJS, Bash, JS/HTML/CSS, Spanner, BigTable, BigQuery, Spring, GraphQL, OOP, MVC, Algos, Git, CoPilot, Jenkins, XLR, Elastic, Jira, AI, ML, Docker, Kubernetes, Kafka, Rest API, GRPC, Grafana. Desirable Support and optimize Distributed Platform technologies including cloud infrastructure, Linux/Unix, containers, APIs, Java-based services, distributed databases, and modern application platforms. Implement and support Cloud/Distributed architectures, modernization initiatives, API enablement, and enterprise integration capabilities. Knowledge of cloud platforms AWS, GCP or general Cloud fundamentals. Data, Integration & Automation Develop and support enterprise integration solutions utilizing APIs, MQ, Connect:Direct, event-driven architectures, and batch and real-time data integration patterns. Utilize relational and NoSQL databases including DB2, PostgreSQL, Redis, and Couchbase to support critical business applications. Automate operational processes, deployments, monitoring, reporting, and remediation activities using Python, Bash, Ansible, Jenkins, and related technologies. Collaborate with engineering teams to adopt scalable automation and self-service capabilities for deployment, monitoring, and operational support. Observability & Continuous Improvement Implement and utilize monitoring, observability, logging, and analytics solutions using tools such as Splunk, OMEGAMON, RMF/SMF, Sysview, MainView, Dynatrace, AppDynamics, ELK, or equivalent technologies. Analyze operational trends, identify opportunities for optimization, and formulate strategic recommendations to improve platform health and engineering effectiveness. Contribute to continuous improvement initiatives focused on reliability, performance, security, operational maturity, and customer experience. DevOps, Security & Governance Understand CI/CD pipelines and DevOps practices using tools such as Git, Jenkins, Maven, DBB, Endevor, Changeman, ISPW, UrbanCode Deploy, or equivalent platforms. Apply enterprise security controls, compliance requirements, audit standards, and access management practices, including RACF, ACF2, Top Secret, and cloud security principles. Ensure solutions meet non-functional requirements (NFRs) including availability, scalability, performance, security, recoverability, and maintainability. Bachelor's degree in Computer Science, Computer Engineering, and/or comparable experience Work experience in software engineering, app support or infrastructure operations or runtime engineering. A working understanding of cloud infrastructure, distributed systems, and containerization technologies, with experience in supporting critical business applications being a plus. Familiarity with monitoring and logging tools, and incident management best practices, to ensure reliability and performance of applications in a production environment. Solid programming and scripting skills, with hands on experience to automate operational tasks using tools such as Python Knowledge of scripting languages (e.g., PowerShell, Python) for automation tasks Experience in technology operations work Hands on experience with relational and NoSQL databases such as DB2, Redis, Postgres, Couchbase etc. Experience in cloud platforms such as AWS, Azure, or Google Cloud, Public Cloud certification is a plus
Specialist, ServiceNow Support Engineering/Operations Support - Technology Workflows
Msd
Specialist, ServiceNow Support Engineering/Operations Support - Technology Workflows
Msd
Network Operations Engineering, Vice President
Blackrock
Manager, Engineering, Git & Gitaly Operations
GitLab
Director, Engineering, Platform Operations & Productivity
GitLab
MT- Remote Korean <> English Interpreters
Intelexoutsourcing