The Night Shift Incident Manager is responsible for leading, coordinating, and overseeing the response to IT incidents that impact the NIH infrastructure, services, and users. This role acts as the central point of authority during incidents, ensuring rapid restoration of service, effective communication, and alignment among technical teams, stakeholders, and leadership. The Incident Manager maintains situational awareness, drives decision‑making, and ensures that incident handling follows established processes and quality standards. This person also, leads a lot of Maintenance Coordination activities because of the hours that are worked. Key Tasks & Responsibilities Major Incident Leadership: Serve as the primary point of contact and authority for all incident‑related decisions. Assess incident severity, declare incident levels (e.g., P1/P2 etc.), and escalate as needed. Quickly evaluate technical information from monitoring tools such as SL1, SiteScope, Sunburst, and network telemetry. Coordinate troubleshooting efforts with internal teams (TOC, network engineering, server teams) and external partners (e.g., Lumen field operations). Communication & Stakeholder Coordination: Direct all communications related to the incident, including updates to leadership, impacted ICs, and service owners. Set up and manage communication bridges when required in MS Teams. Keep internal and external stakeholders informed with accurate, timely, and concise updates. Actively mitigate panic, confusion, or misalignment through calm, structured messaging. Technical Oversight & Coordination: Ensure all required technical SMEs are engaged and aligned on priorities. Request additional resources, expertise, or escalation paths when necessary. Maintain awareness of all activities performed during the incident and ensure efficient workflow. Confirm the validity of troubleshooting steps and ensure teams are not duplicating work. Decision‑Making & Delegation: Make confident and timely decisions based on available data and SME input. Delegate tasks to team members and expand the team as incident complexity increases. Maintain focus on the broader incident timeline while individual teams address specific areas. Post‑Incident Activities: Lead and document post‑mortem reviews; ensure lessons learned are captured. Produce incident documentation including incident reports, PRBs, ITASK records, and communication summaries. Recommend process improvements, preventative actions, and training needs. On‑Call Responsibilities: (Reflecting CIT On‑Call Schedule expectations) Provide on‑call coverage during assigned rotation periods on the weekends. You are not required to come on site but work remotely if there is a problem. Always be reachable with phone and laptop available. Respond within required timeframes (e.g., within 15 minutes); delays may result in corrective action. Notify supervisors of conflicts or inability to cover assigned rotations. Supervisory Responsibilities: Approve Leave and Timesheets. Mentor and Coach other TOC Engineers especially when they are having trouble with a process or procedure. Notify Leadership of any problems or situations that may arise. Provide guidance to the TOC Engineers by engaging in nightly operational activities. Education & Experience Required Skills & Competencies: Bachelor’s degree from an 4 year accredited school program and 4 years of related experience. Associates degree from an accredited program and 6 years of related experience. Strong verbal and written communication skills. High‑level understanding of NIH/CIT infrastructure, monitoring tools, and ITSM processes. Ability to synthesize complex information quickly and make informed decisions. Calm and steady demeanor under pressure; ability to manage stressed personnel. Strong leadership qualities, including ability to direct diverse technical teams. Proven problem‑solving skills and analytical thinking. Experience participating in or leading major incident responses. Preferred Experience Prior experience with the ITIL Framework and certification. Prior experience in incident management or similar operational command roles. Familiarity with NIH network topology, data centers, and vendor support escalation. Experience with ITSM platforms (ServiceNow), monitoring tools (SL1, SiteScope), and communication workflows such as xMatters. Participation in mock drills, communication workshops, or incident management training. Work Environment Both a fast-paced and slow-paced operational setting, because of the off hours, involving real-time troubleshooting and cross-team coordination when possible. Most night events carry over to the morning to our Day Shift. Frequent communication across OMS, network engineering teams, ICs, and external vendors. Participation in TOC meetings, weekly incident reviews, and documentation improvement discussions. Security Clearance Must be able to obtain government customer clearance. Other (Travel, Work Environment, DoD 8570 Requirements, Administrative Notes, etc.) Normal Shift Hours 10PM-6:30AM EST M-F This team provides 24x7x365 support to the user community. Program covers a 24/7 operation, and members are asked to be flexible in providing coverage outside of their normal shift hours, when the need arises. This is also a Hybrid Position where you would come into work certain days on-site in Bethesda, MD and teleworking some days from home. Computer World Services is an affirmative action and equal employment opportunity employer. Current employees and/or qualified applicants will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, disability, protected veteran status, genetic information or any other characteristic protected by local, state, or federal laws, rules, or regulations. Computer World Services is committed to the full inclusion of all qualified individuals. As part of this commitment, Computer World Services will ensure that individuals with disabilities (IWD) are provided reasonable accommodations. If reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please contact Human Resources at [email protected] .
Product Marketing Manager - Incident Response
Datadog
Incident Manager (Global Incident Management) - Assistant Vice President
Db
Senior Manager, Incident Management
Zillow
CFC Senior Enterprise Security Incident Manager
Experian
Senior Incident Manager
Snowflake
US Program Manager, Cyber Monitoring & Incident Response
DTCC Candidate Experience Site