Top Skills' Details - Experience designing and implementing enterprise observability solutions across metrics, logs, and distributed traces. - Experience developing instrumentation standards using OpenTelemetry and modern observability frameworks. - Experience building dashboards, alerts, and service health indicators that support application reliability and business outcomes. Description Our client is seeking a Senior Observability Engineer to help drive the evolution of enterprise observability capabilities across a large-scale cloud environment. This individual contributor will play a key role in developing observability standards, improving platform visibility, and enabling engineering teams to proactively monitor, troubleshoot, and optimize critical applications and services. This position is ideal for someone who is passionate about monitoring, telemetry, distributed systems, and reliability engineering and enjoys partnering with development, SRE, and platform teams to improve operational excellence at scale. Key Responsibilities: Observability Engineering Design and implement enterprise observability solutions across metrics, logs, and distributed traces. Develop instrumentation standards using OpenTelemetry and modern observability frameworks. Build dashboards, alerts, and service health indicators that support application reliability and business outcomes. Improve visibility across development, test, and production environments. Partner with engineering teams to identify and close monitoring gaps. Platform Optimization: Support ongoing consolidation and modernization of observability tooling. Develop strategies to improve signal quality while reducing telemetry noise. Help establish governance standards for monitoring, alerting, and data ingestion. Drive best practices around observability adoption across multiple engineering organizations. Recommend improvements to platform reliability, scalability, and performance. Reliability & Operations: Support incident investigation through effective monitoring and root-cause analysis. Develop service-level indicators (SLIs), service-level objectives (SLOs), and reliability metrics. Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through enhanced observability practices. Participate in architectural discussions around resiliency and operational readiness. Cost Optimization: Analyze telemetry consumption and platform usage trends. Implement strategies for reducing unnecessary ingestion costs. Help standardize logging, retention, sampling, and observability governance practices. Balance visibility requirements with overall platform efficiency. Required Qualifications 5+ years of experience in Observability Engineering, SRE, DevOps, Platform Engineering, or related disciplines. Strong knowledge of: Metrics Logs Distributed Tracing Alerting & Monitoring Strategies Hands-on experience with OpenTelemetry and Prometheus. Experience supporting cloud-native environments (AWS preferred). Experience with observability platforms such as: Coralogix Datadog Splunk Grafana New Relic Understanding of microservices architectures and cloud-based systems. Experience troubleshooting production environments using telemetry data. Strong communication and collaboration skills. Preferred Qualifications Experience with Terraform or Infrastructure as Code. Exposure to CI/CD platforms including GitHub, Jenkins, or Azure DevOps. Knowledge of SRE principles and operational excellence frameworks. Experience supporting enterprise-scale observability initiatives. Familiarity with platform engineering and developer enablement practices. What Success Looks Like Improved observability coverage across critical systems. Reduced monitoring gaps between non-production and production environments. Lower telemetry costs while maintaining platform visibility. Faster incident detection and resolution. Increased adoption of observability standards across engineering teams. Skills Linux, Cloud, Python, Aws, Devops, Automation, Azure, Administration, Active directory, Terraform, Kubernetes, Security Top Skills Details Linux,Cloud,Python,Aws,Devops,Automation,Azure,Administration,Active directory,Terraform,Kubernetes,Security Additional Skills & Qualifications N/A Experience Level Expert Level Job Type & Location This is a Contract to Hire position based out of Raleigh, NC. Pay and Benefits The pay range for this position is $70.00 - $85.00/hr. Individual compensation offered for this position within this range will depend on many factors, including qualifications, skills, relevant experience, job knowledge, geographic location, internal equity, and other pertinent job-related factors. Eligibility requirements apply to some benefits and may depend on your job classification and length of employment. Benefits are subject to change and may be subject to specific elections, plan, or program terms. If eligible, the benefits available for this temporary role may include the following: • Medical, dental & vision • Critical Illness, Accident, and Hospital • 401(k) Retirement Plan – Pre-tax and Roth post-tax contributions available • Life Insurance (Voluntary Life & AD&D for the employee and dependents) • Short and long-term disability • Health Spending Account (HSA) • Transportation benefits • Employee Assistance Program • Time Off/Leave (PTO, Vacation or Sick Leave) Workplace Type This is a fully remote position. Application Deadline This position is anticipated to close on Aug 26, 2026. About TEKsystems We're partners in transformation. We help clients activate ideas and solutions to take advantage of a new world of opportunity. We are a team of 80,000 strong, working with over 6,000 clients, including 80% of the Fortune 500, across North America, Europe and Asia. As an industry leader in Full-Stack Technology Services, Talent Services, and real-world application, we work with progressive leaders to drive change. That's the power of true partnership. TEKsystems is an Allegis Group company. The company is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law. About TEKsystems and TEKsystems Global Services We’re a leading provider of business and technology services. We accelerate business transformation for our customers. Our expertise in strategy, design, execution and operations unlocks business value through a range of solutions. We’re a team of 80,000 strong, working with over 6,000 customers, including 80% of the Fortune 500 across North America, Europe and Asia, who partner with us for our scale, full-stack capabilities and speed. We’re strategic thinkers, hands-on collaborators, helping customers capitalize on change and master the momentum of technology. We’re building tomorrow by delivering business outcomes and making positive impacts in our global communities. TEKsystems and TEKsystems Global Services are Allegis Group companies. Learn more at TEKsystems.com. The company is an equal opportunity employer and will consider all applications without regard to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law. San Francisco Fair Chance Ordinance: Pursuant to the San Francisco Fair Chance Ordinance, for all positions located in the city and county of San Francisco, we will consider for employment qualified applicants with arrest and conviction records. Massachusetts Lie Detector: It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability. Use of Artificial Intelligence (AI): We may use Artificial Intelligence (AI) to support parts of our hiring process, including sourcing, screening, and evaluating candidates. AI helps assess applications and qualifications, but final decisions are made by our hiring team. By applying, you acknowledge and agree that your application may be reviewed using AI tools.
Lead Product Engineer – (AIOps for Observability) - Remote
Allstate
Sr Engineer - Applications & Observability
Unisys
Observability & SRE Engineer
Wwtsoftchoice
Staff Software Engineer , Observability
Observability & SRE Engineer
Wwt
Site Observability Engineer
Bright Vision Technologies