AIOps & Observability Lead
- Salary
- $140K–$180KUSD per year
- Hiring from
- United States
- Work type
- Hybrid
- Posted
- Sep 25, 2026
Role Summary
We are building a new Operations function and need a hands-on, technical leader to set the direction for our AIOps and observability strategy. This is a player-coach role, weighted toward player: someone who works directly in the tools, not just the roadmap.
Our new Splunk Observability implementation is a foundation of a broader automated operational intelligence ecosystem. We will lean on managed-service partners to act as first-line "eyes on glass" for the alerts and events it generates — but this role owns the strategy, architecture, and quality behind what those partners are watching. The goal is not just detection; it's ensuring issues can be discovered, triaged, and resolved through tooling and automation, reducing L3 engineer escalations wherever possible.
The role could expand over time into desktop/endpoint observability, including a Nexthink implementation.
Key Responsibilities
- Define and own the observability and event management strategy, stack, and processes, including how major incidents are detected, triaged, and resolved.
- Lead the rollout of Splunk Observability across our application and infrastructure portfolio, driving onboarding, instrumentation, and coverage expansion.
- Lead implementation of event management tooling — such as, BigPanda, Splunk ITSI, or PagerDuty — to correlate, prioritize, and route alerts from Splunk Observability.
- Coach and enable the Automation team to become proficient in the observability and event intelligence platform, including alert logic, dashboard creation, event enrichment, correlation workflows, triggered automations, MCP-style integrations, and runbook automation.
- Partner with the Infrastructure Architect to build a roadmap toward first-class, enterprise-grade observability.
- Develop and execute an AIOps strategy that turns telemetry into HITL-automated, self-healing operations — evaluating and building on tooling in Splunk, AWS, or elsewhere as appropriate.
- Stand up and mature an AIOps/SRE capability: automated remediation, intelligent alert correlation, and reduced manual triage.
- Direct managed-service partners providing "eyes on glass" monitoring — defining what they watch, how they escalate, and holding them accountable to SLAs.
- Design escalation paths so that issues resolvable via tooling are handled there first, protecting L3 engineering time for what truly needs it.
- Establish standards for event classification, severity, ownership, suppression, and closure across the environment.
- Assess and help scope the expansion into desktop/endpoint observability, including a Nexthink implementation.
- Ensure dashboards, alerts, and event workflows are built around operational decisions, not just technical visibility.
- Align observability, event management, incident, problem, change, knowledge, and escalation workflows with ITIL/ITSM practices and ServiceNow processes.
Required Qualifications
- Strong, hands-on experience with observability, monitoring, and event management platforms (e.g., Splunk, Splunk ITSI, BigPanda, PagerDuty).
- Proven experience designing and implementing AIOps or event-intelligence capabilities — alert correlation, noise reduction, and automation.
- Experience defining incident management processes, including major incident response and escalation design.
- Comfortable operating as both strategic lead and hands-on technical contributor — able to shape direction and work directly in the tools.
- Scripting or automation experience (e.g., Python, PowerShell, REST APIs) to support event enrichment and automated remediation.
- Strong communication skills, with the ability to align infrastructure, application, and operations teams around a shared observability strategy.
Preferred Qualifications
- Experience with AWS or other cloud-native observability and automation tooling.
- Experience with endpoint/desktop observability platforms such as Nexthink.
- ITIL/ITSM familiarity, particularly incident, problem, and change management.
- Experience managing or directing managed-service/outsourced monitoring partners.
- Experience in regulated, high-availability, or large-scale enterprise environments.
Success Profile
The right person has been in the room during major incidents, knows what a noisy, low-trust alerting environment looks like, and knows what it takes to fix it. They can set a multi-year observability direction and also configure the tool themselves when needed. Success looks like fewer incidents reaching L3, faster time-to-resolution, and an Operations function that trusts its own data.
The base salary range for this position is $140,000 - $180,000 per year. This range reflects the minimum and maximum base salary we reasonably expect to pay for this role. In addition, this position may be eligible to participate in the relevant business unit’s incentive compensation plan, and other compensation programs as applicable. Eligible employees may participate in a 401(k) program with a generous profit-sharing contribution, medical, prescription dental, and vision coverage; life insurance; disability coverage; paid holidays; vacation; and sick time, subject to plan terms and Company policies.
About Bessemer Trust:
- Bessemer Trust is a family office, overseeing $250 billion in assets for over 3,000 individuals and families of substantial wealth. Its more than 1,300 employees are singularly focused on private wealth management — disciplined investment management, sophisticated wealth planning, comprehensive family office services, and highly personalized client service.
- Established in 1907 as the family office for Annie and Henry Phipps, Bessemer Trust is in its seventh generation of ownership by the Phipps family. As a self-made entrepreneur, Henry Phipps was a founding partner and chief financial officer of Carnegie Steel.
- Bessemer Trust retains its original focus as a privately owned and independent wealth manager deeply committed to its mission of providing peace of mind to its clients. Bessemer’s adherence to putting clients’ interests first, fiduciary mindset, and highly collaborative culture are at the heart of everything the firm does.
Key Facts:
- For more than 119 years, Bessemer Trust has operated continuously in a single line of business, independently owned by one family.
- Headquartered in New York’s Rockefeller Center, Bessemer Trust has 22 offices in total. Woodbridge, NJ, is one of the firm’s largest offices, which hosts a wide range of technology and operations professionals. In addition to its sizable presence in New York and Woodbridge, the firm provides client service through offices in Atlanta, Boston, Chicago, Dallas, Delaware, Denver, Garden City, Grand Cayman, Greenwich, Houston, Los Angeles, Miami, Naples, Nevada, Palm Beach, San Diego, San Francisco, Seattle, Stuart, and Washington, D.C.
- To watch a video about Bessemer Trust’s history, click here.
- To learn more about Bessemer Trust, click here.
About Our Employee Rewards and Benefits:
- We provide exceptional rewards and benefits that are among the best in the industry, giving our people access to a wide range of options, including:
- Competitive base salary plus discretionary annual bonus for select positions
- A 401(k) plan with a generous annual profit-sharing contribution
- Personalized development and career opportunities, including tuition reimbursement support
- Comprehensive medical, dental, and vision plans with zero contributions for employee coverage
- Employee assistance (EAP) and wellness programs
- Hybrid work environment: 60% in office, 40% remote for most positions
- Paid time off and paid parental leave
- Employer-paid life insurance and short- and long-term disability coverage
- Legal services and financial wellness plans at no cost to employees
Bessemer Trust is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. We encourage candidates of diverse backgrounds to apply.