Key Responsibilities Ensure 99.99% or greater infrastructure availability through proactive monitoring, maintenance coordination, operational controls, and disciplined incident management. Oversee the health of infrastructure and services, including compute, storage, memory, network, cloud, database, application, endpoint, and integration components. Drive automation-first operations through self-healing pipelines, automated remediation playbooks, and infrastructure-as-code patterns that reduce manual toil, standardize recovery actions, and improve mean time to restore service. Ensure alerts are acknowledged promptly, validated against known conditions, correctly categorized, prioritized, documented, and routed to the appropriate resolver group. Review alert thresholds, suppression rules, correlation logic, maintenance windows, and routing policies to improve signal quality and reduce avoidable noise. Direct initial troubleshooting using approved runbooks, knowledge articles, dashboards, logs, and diagnostic tools. Coordinate rapid escalation of warning and exception events that indicate service degradation, capacity risk, or potential incident. Serve as the primary escalation point for operational incidents, lead root cause analysis, and drive corrective and preventive actions through completion. Support major incident response by establishing situational awareness, assigning monitoring actions, maintaining an event timeline, and providing accurate technical updates. Analyze recurring alerts, resource trends, capacity indicators, service dependencies, and monitoring gaps; initiate corrective actions with engineering and problem management teams. Own Operating procedures, runbooks, escalation matrices, contact lists, shift checklists, and knowledge documentation. Produce operational summaries and periodic reports covering service health, significant events, response performance, recurring conditions, and improvement actions. Maintain clear shift coverage, handoff, attendance, workload, and escalation expectations across the monitoring function. Act as the operational bridge among monitoring engineers, incident management, service owners, application teams, infrastructure teams, vendors, and business stakeholders. Lead post-event reviews and continual improvement initiatives focused on automation, observability, runbook quality, engineer readiness, and faster restoration. Required Qualifications Seven or more years of experience in IT operations, infrastructure support, network operations, cloud operations, application support, site reliability, or a related field. 2 years of experience leading a Network Operations team or in a team lead capacity. Preferred. Strong working knowledge of ITIL monitoring and event management, incident management, problem management, change enablement, configuration management, and service-level management. Hands-on experience with enterprise monitoring, observability, alerting, ticketing, paging, log analysis, and dashboard platforms. Hands-on experience integrating Grafana, PagerDuty, ThousandEyes, Azure Monitor, Amazon CloudWatch, or comparable tools, within an enterprise observability ecosystem. Hands-on experience troubleshooting Windows Server operating systems, services, performance, connectivity, storage, patching, and related infrastructure issues in production environments. Strong knowledge of AWS and Microsoft Azure platforms, with hands-on experience troubleshooting cloud-hosted services. Ability to interpret infrastructure and application telemetry, identify service impact, prioritize competing conditions, and make timely decisions under pressure. Demonstrated experience managing shifts, on-call coverage, escalations, runbooks, operational metrics, and cross-functional response. Clear written and verbal communication skills, including the ability to translate technical conditions into concise business-impact updates. Strong leadership, coaching, collaboration, organization, and continuous-improvement skills. Benefits Veradigm believes in empowering our associates with the tools and flexibility to bring the best version of themselves to work. Through our generous benefits package with an emphasis on work/life balance, we give our employees the opportunity to allow their careers to flourish. Quarterly Company-Wide Recharge Days Flexible Work Environment (Hybrid) Peer-based incentive “Cheer” awards Tuition Reimbursement Program To know more about the benefits and culture at Veradigm, please visit the links mentioned below: - https://veradigm.com/about-veradigm/careers/benefits/ https://veradigm.com/about-veradigm/careers/culture/ #LI-SL1 #LI-Hybrid Veradigm is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse and inclusive workforce. Thank you for reviewing this opportunity! Does this look like a great match for your skill set? If so, please scroll down and tell us more about yourself!
Network Operation Engineer
Share
Operations Business Specialist, ESC
Genmills
Consultant - Operational Technology Data Operations
Genmills
Senior Manager -Tech Ops Engineering
American Express
Technical Program Manager III
Ffive
Senior Executive
EXL Talent Acquisition Team