SI

Monitoring & Observability Engineer

Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 27, 2026
Is this job info correct?
We are looking for a Monitoring & Observability Engineer to support the design and build-out of a modern observability and event management capability, with a strong focus on Datadog, ServiceNow integration, monitoring policy migration, and event pipeline modernization.

This role will be responsible for extracting, analyzing, and re-mapping existing BMC TrueSight/Helix monitoring policies and event rules into a new observability model. The ideal candidate has hands-on experience building or significantly improving monitoring and event management environments, not just maintaining existing tools.

This person will play a key role in designing log and metric pipelines, supporting SNMP trap processing, implementing Datadog-ServiceNow integration, and helping define a tiered observability strategy that balances production-critical visibility with cost-efficient lower-tier monitoring using OpenTelemetry-based open-source pipelines.

Key responsibilities:
  • Analyze existing BMC TrueSight/Helix monitoring policies, alerts, thresholds, and event rules
  • Re-map legacy monitoring logic into Datadog and supporting observability pipelines
  • Design and configure log, metric, event, and alert pipelines
  • Support SNMP trap processing, normalization, enrichment, and routing
  • Stand up and configure net-new Datadog-ServiceNow integration
  • Define event correlation, deduplication, escalation, and incident creation logic
  • Support migration from legacy monitoring platforms to Datadog and OpenTelemetry-based pipelines
  • Help design a tiered observability model for production, non-production, and lower-tier workloads
  • Use open-source/OpenTelemetry pipelines where appropriate to optimize Datadog usage and cost
  • Collaborate with infrastructure, network, application, and ITSM teams to align monitoring coverage with operational needs
  • Document monitoring standards, event rules, integration patterns, and operational procedures
  • Support testing, validation, and tuning of monitoring rules and event flows
Required qualifications:
  • Bachelor’s degree in Information Technology, Computer Science, Engineering, or related field, or equivalent practical experience
  • 5–7 years of experience in monitoring, observability, event management, or IT operations engineering
  • Hands-on experience with Datadog or similar observability platforms
  • Experience with BMC TrueSight, BMC Helix, or comparable enterprise monitoring/event management tools
  • Strong understanding of logs, metrics, alerts, events, thresholds, and incident workflows
  • Experience designing or supporting monitoring/event pipelines
  • Experience with SNMP traps, network/device events, and event normalization
  • Familiarity with ITSM platforms, especially ServiceNow
  • Understanding of incident management, alert routing, escalation, and operational support models
  • Ability to analyze existing monitoring rules and translate them into a modern observability framework
  • Strong troubleshooting, documentation, and communication skills
Preferred qualifications:
  • Experience building an event management capability from the ground up
  • Experience with Open Telemetry collectors, agents, or open-source observability pipelines
  • Experience with Datadog-ServiceNow integration
  • Familiarity with cost optimization strategies for observability platforms
  • Experience supporting enterprise infrastructure, network, cloud, or hybrid environments
  • Knowledge of event correlation, alert noise reduction, and monitoring rationalization
  • Experience working in transformation, migration, or platform modernization programs
What success looks like in this role:
  • Legacy BMC TrueSight/Helix monitoring policies are clearly analyzed, documented, and mapped to target-state observability platforms
  • Datadog is configured to support production-critical monitoring with reliable event and incident flows
  • ServiceNow receives clean, actionable, and properly routed events from Datadog
  • SNMP trap processing is stable, normalized, and aligned with operational needs
  • Lower-tier and non-production workloads are monitored through cost-effective OpenTelemetry-based pipelines
  • Alert noise is reduced, monitoring coverage is improved, and support teams have clearer visibility into incidents
  • Observability standards and documentation are in place for future operational support
What we offer:
  • Competitive pay and salary growth based on your performance.
  • Remote setup, B2B long-term contract, US EST working hours
  • Dynamic working culture with opportunities to grow.
  • Challenging projects with cutting-edge AI and cloud technologies.
  • 100% paid sick leave.
  • Friendly colleagues and a pleasant working atmosphere.
Thank you for your interest in this position. Please note that only candidates whose qualifications closely match our requirements will be contacted.

Similar jobs

Apply for this job