Are you ready to make an impact at DTCC? Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve. The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance. Pay and Benefits: Competitive compensation, including base pay and annual incentive Comprehensive health and life insurance and well-being benefits Pension Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being. DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee). The impact you will have in this role The Senior Principal, Operations Resiliency & GameDay Strategy within Cloud Operations is responsible for defining and operating the enterprise GameDay program that validates the resiliency of cloud-hosted applications and platforms. You will design failure scenarios drawn from real production incidents, select and prioritize applications for testing, and ensure that exercise outcomes translate directly into improved runbooks, recovery paths, and architectural resilience. Working across Cloud Operations, Incident Management, and application engineering teams, you will drive a closed-loop process where what breaks in production is systematically tested, validated, and retested until recovery capabilities are proven and measurable. Your primary responsibilities Define and operationalize the enterprise GameDay strategy that validates resiliency of cloud-hosted applications and platforms. Establish application selection criteria, testing frequency, and a scalable operating model for controlled failure testing. Develop a taxonomy of application patterns and map them to targeted, reusable fault injection scenarios drawn from real production failures. Oversee GameDay execution, manual and automated, validating observability, detection, response, and recovery workflows. Close the loop between production incidents and simulated testing: incident, scenario, test, finding, mitigation. retest. Drive identification of missing recovery paths, observability gaps, and ineffective automation, then ensure findings are tracked and assigned to accountable owners. Push execution from manual to automated: fault injection, scenario orchestration, metrics capture, and reporting. Define and track effectiveness metrics: time to detect, time to recover, automated recovery success rates, and reduction in repeat incident patterns. Partner across Incident Management, Application Teams, Platform Engineering, and SRE to ensure GameDay outputs translate directly into improved runbooks, recovery paths, and architectural resilience. Qualifications: Minimum of 10 years of related experience Bachelor's degree preferred or equivalent experience Talents Needed for Success: Deep experience in cloud operations, SRE, or resiliency engineering Hands-on knowledge of incident management and root cause analysis Experience with fault injection tools such as AWS FIS, Gremlin, or equivalent Experience with AI-assisted operations tooling such as AWS DevOps Agent or equivalent Familiarity with policy-as-code frameworks such as OPA, Sentinel, or equivalent Strong understanding of distributed systems failure modes, observability architectures, and automation/runbook engineering Ability to operate horizontally across engineering and operations organizations — influencing without direct delivery ownership Strong judgment on what to test, when, and why We offer top class training and development for you to be an asset in our organization! We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
Mechanical Design Engineer - Product Development (Remote)
Jobs Ai
Software Development Engineer - macOS Endpoint
BeyondTrust
Software Development Engineer
Adobe
Software Development Engineer
Adobe
Senior Business Development Manager - Engineering Biology
LGC Group
Head of Business Development - Lotus Engineering
Uslot