Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Nscale logo

Operational Data & Observability Engineer

Nscale
Posted 3 weeks ago
🇺🇸United States🏢Hybrid📁Engineering & Development
Is this job info correct?

. Operational Data & Observability Engineer About the Role We're looking for an Operational Data & Observability Engineer to build and evolve the monitoring, logging, and observability capabilities that power our production environments. In this role, you'll help ensure our infrastructure and applications remain reliable, scalable, and performant by providing engineering teams with actionable operational insights. You'll partner closely with DevOps, Site Reliability Engineering (SRE), platform, and software engineering teams to develop modern observability solutions, improve incident response, and enable data-driven operational excellence. What You'll Do Design & Build Observability Solutions Design and implement enterprise observability strategies across infrastructure, services, and applications. Develop monitoring dashboards, alerts, and Service Level Objectives (SLOs) that provide meaningful operational visibility. Build and maintain centralized logging and log analysis pipelines. Implement distributed tracing to improve visibility across microservices and complex application workflows. Establish performance baselines and develop anomaly detection strategies. Operational Data Engineering Deploy, configure, and maintain metrics, logs, events, and telemetry collection systems. Design and manage operational data pipelines that support monitoring and analytics. Develop APIs and integrations that enable operational data consumption across teams. Ensure data quality, consistency, retention, and cost-efficient storage practices. Reliability & Operations Troubleshoot production issues using monitoring, logging, and tracing data. Participate in an on-call rotation and support incident response activities. Create and maintain operational documentation, runbooks, and troubleshooting guides. Partner with engineering teams to improve platform reliability, scalability, and operational readiness. Continuously optimize observability infrastructure for performance and resilience. Platform & Tool Administration Administer and enhance observability platforms such as Datadog, Grafana, Prometheus, ELK Stack, New Relic, or similar technologies. Evaluate emerging observability tools and recommend improvements. Automate monitoring deployments, instrumentation, and platform configuration. Perform ongoing maintenance, upgrades, and lifecycle management of observability infrastructure. What You'll Bring Required Qualifications 3+ years of experience in DevOps, Site Reliability Engineering (SRE), Operations Engineering, Platform Engineering, or Observability Engineering. Hands-on experience with modern monitoring platforms such as Prometheus, Grafana, Datadog, New Relic, or equivalent. Experience working with centralized logging platforms including ELK/Elastic Stack, Splunk, CloudWatch, or similar solutions. Proficiency with scripting or programming languages such as Python, Go, Bash, or equivalent. Strong understanding of observability fundamentals, including metrics, logging, distributed tracing, and application performance monitoring (APM). Experience working with cloud platforms (AWS, Azure, or Google Cloud Platform) and Kubernetes or other container orchestration technologies. Solid understanding of application, infrastructure, networking, database, and storage performance monitoring. Strong analytical, troubleshooting, communication, and documentation skills with a collaborative approach to problem-solving. Preferred Qualifications Experience supporting microservices-based architectures. Expertise across multiple observability platforms. Experience with incident management, root cause analysis, and post-incident reviews. Infrastructure as Code experience using Terraform, Ansible, or similar tools. Familiarity with eBPF or low-level Linux performance monitoring. Experience building custom telemetry, ETL, or operational data pipelines. Understanding of security monitoring, audit logging, and compliance requirements. What Success Looks Like Success in this role will be measured by your ability to: Improve platform visibility and operational health. Reduce Mean Time to Resolution (MTTR) during incidents. Increase alert quality while reducing unnecessary noise. Deliver highly available, scalable observability platforms. Improve engineering productivity through actionable monitoring and operational insights. Optimize observability infrastructure performance and cost efficiency. Work Environment Participate in a rotating on-call schedule to support production environments. Support mission-critical systems with occasional after-hours or incident response responsibilities. Hybrid or remote work arrangements available, depending on business needs. Why Join Us? You'll play a critical role in building the operational intelligence that keeps our platforms running at scale. If you're passionate about observability, automation, reliability, and empowering engineering teams with meaningful operational insights, we'd love to hear from you. The range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation. Salary Range $145,000 — $180,000 USD For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.

Similar jobs

Similar jobs

Grafana Labs logo

Senior FullStack Engineer - Grafana Cloud Observability| US | Remote

Grafana Labs

🇺🇸United StatesYesterday
Bright Vision Technologies logo

Observability Engineer

Bright Vision Technologies

🇺🇸United States4 days ago
FreedomPay logo

Sr. Observability Engineer

FreedomPay

🇺🇸United States5 days ago
Honeycomb.io logo

Senior Software Engineer II - LLM Observability

Honeycomb.io

🇺🇸United States1 weeks ago
ECS Federal LLC logo

Senior Kibana/Observability Engineer

ECS Federal LLC

🇺🇸United States1 weeks ago
Datadog logo

Director, Engineering - Cloud Observability

Datadog

🇺🇸United States2 weeks ago