Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
MeridianLink logo

Sr. Site Reliability Engineer

MeridianLink
Posted 1 hour ago
🇺🇸United States🏠Remote💰$104.1K–$177.6K📁Engineering & Development
Is this job info correct?

About the Role We are seeking a Senior Site Reliability Engineer to join our cloud engineering team. You will own the reliability, scalability, and observability of our critical financial SaaS applications and infrastructure, working across cloud platforms to ensure our customers experience is seamless, secure, and performant services. This is a high-impact role for someone who is passionate about building resilient systems and preventing outages before they happen. Key Responsibilities Design, implement, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across all critical systems; ensure we meet or exceed targets consistently Lead observability strategy by designing comprehensive monitoring, logging, and tracing architectures; select and deploy observability tools that provide deep visibility into system behavior Build and own runbooks, incident response procedures, and post-incident review processes; mentor the team on incident management and blameless postmortems Architect and deploy cloud infrastructure on AWS or Azure; implement infrastructure-as-code practices and ensure high availability, disaster recovery, and business continuity Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable self-healing systems; implement intelligent alerting to minimize false positives Drive reliability improvements through load testing, chaos engineering, and failure scenario analysis; identify and eliminate single points of failure Partner with application and backend teams to design reliable systems from inception; conduct architecture reviews and reliability assessments Write production-grade Python tooling for automation, metrics collection, alert management, and operational workflows Champion security and compliance in infrastructure; implement defense-in-depth principles for a regulated fintech environment Required Qualifications 7+ years in Site Reliability Engineering, DevOps, platform engineering, or closely related roles with significant responsibility for production systems Expert-level experience with Azure or AWS (or both); deep knowledge of compute, networking, storage, and managed services; experience managing infrastructure at scale Demonstrated expertise in observability: designing and implementing monitoring, alerting, logging, and distributed tracing solutions; hands-on with observability platforms (e.g., Prometheus, Grafana, ELK, Datadog, New Relic, or similar) Strong background in SLOs, SLIs, and SLAs; experience defining meaningful objectives and building systems to meet them; understanding of error budgets and their role in prioritization Proven experience designing and troubleshooting highly available, resilient, and scalable systems; deep understanding of distributed systems concepts and failure modes Proficiency in Python, PowerShell, bash, etc. scripting languages for production automation, tooling, and systems programming; ability to write clean, maintainable code for operational workflows Hands-on experience with AIOps practices: event correlation, intelligent alerting, predictive analytics, and automated remediation; familiarity with AIOps platforms is a plus Experience with infrastructure-as-code tools (e.g., Terraform, CloudFormation, Ansible); version control and CI/CD pipeline design Track record of incident management and on-call ownership; comfort with incident response and the ability to remain calm under pressure Excellent communication skills; ability to work cross-functionally and influence without authority; comfort mentoring junior engineers Preferred Qualifications Experience in the fintech, payments, banking, or other regulated industries; understanding of compliance requirements (SOC 2, PCI-DSS, etc.) Experience with Kubernetes and container orchestration; deep knowledge of containerized application deployment and management Proficiency with observability as code; experience building custom metrics, dashboards, and alerts programmatically Background in chaos engineering or reliability testing; experience using tools like Gremlin or similar platforms Contribution to open-source observability or infrastructure projects Expertise in network security, application security, or infrastructure hardening Experience with database optimization, query performance tuning, and backup/recovery strategies

Similar jobs

Similar jobs

BeyondTrust logo

Staff Site Reliability Engineer

BeyondTrust

🌍Canada, United States2 hours ago
Circle logo

Staff Site Reliability Engineer

Circle

🇺🇸United States18 hours ago
Treeswift logo

Senior Site Reliability and Infrastructure Engineer

Treeswift

🇺🇸United States19 hours ago
Nutanix logo

Systems Reliability Engineer II

Nutanix

🌍India, United States19 hours ago
Careers Inc logo

Site Reliability Engineer

Careers Inc

🌍Mexico, United States20 hours ago
Circle logo

Staff Site Reliability Engineer

Circle

🇺🇸United States20 hours ago