Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Heirs Technologies logo

SRE Lead

Heirs Technologies
Posted 12 hours ago
🌍Worldwide🏠Remote📁Engineering & Development
Is this job info correct?

About the Engagement We are building a next-generation digital banking platform for one of Africa's largest financial institutions --- a product built to serve 15 million users at launch, scaling to hundreds of millions across the continent and diaspora. On a financial platform at this scale, reliability is not a feature --- it is a promise. This is a once-in-a-generation opportunity to build the reliability function that keeps that promise. The Role The SRE Lead is the owner of the platform's reliability function --- the person who defines what reliability means on this platform, builds the systems and processes to deliver it, and leads the team that keeps it. Reporting to the Platform Manager and leading a team of Site Reliability Engineers, you will design and operate the platform's observability stack, incident management framework, and NSOC --- ensuring the platform is always visible, always monitored, and always recoverable. On a banking platform serving millions of Africans, downtime is not just a technical failure --- it is a failure of trust. You are the person who makes sure that never happens. What You'll Do SRE Strategy & Reliability Standards --- Own the platform's Site Reliability Engineering function --- defining the reliability philosophy, standards, and practices that govern how the platform is operated. Establish and enforce SLOs and error budgets for all platform services, working with engineering teams to ensure reliability targets are defined, measured, and taken seriously. Build a culture where reliability is a shared engineering responsibility, not just an ops concern. NSOC Design & Operations --- Design, build, and operate the platform's 24/7 Network & Security Operations Centre (NSOC) --- the central nervous system of the platform's operational awareness. Define the monitoring coverage, alert routing, escalation paths, and on-call workflows that ensure no incident goes undetected and no alert goes unacted upon. The NSOC must be staffed, tooled, and operating continuously from the moment the platform goes live. Observability Architecture & Implementation --- Own the platform's observability stack --- designing and implementing the centralised logging, metrics collection, distributed tracing, and alerting infrastructure that provides complete visibility into every layer of the platform. Ensure every service is instrumented from day one in production, alerting is tuned to minimise noise without missing signal, and dashboards give the team the operational intelligence they need at a glance. Incident Management --- Own the platform's end-to-end incident management process --- from detection and triage through containment, resolution, and post-incident review. Define the on-call framework, escalation paths, and severity classification standards. Ensure every significant incident results in a blameless post-mortem with clear action items tracked to completion. Continuously improve incident response speed and quality --- tracking MTTR as a primary platform health metric. Reliability Engineering & Chaos Testing --- Design and run a proactive reliability engineering programme --- including chaos engineering experiments, load and stress testing, failure mode analysis, and disaster recovery drills. Do not wait for production to reveal failure modes. Identify and address weaknesses in the platform's resilience before they become incidents, and validate that the platform's recovery capabilities work exactly as designed under real conditions. SLO & DORA Metrics --- Own the collection, baselining, and reporting of the platform's reliability and engineering performance metrics --- including SLO attainment, error budget consumption, deployment frequency, lead time for changes, change failure rate, and mean time to recovery (MTTR). Surface these metrics to the Platform Manager in real time and use them to identify where engineering practice and platform reliability need focused improvement. Toil Reduction & Automation --- Systematically identify, measure, and eliminate toil --- the repetitive, manual, automatable operational work that consumes engineering capacity without improving reliability. Build automation that makes the platform easier to operate at scale and frees the SRE team to focus on reliability engineering rather than routine operations. Toil reduction is an ongoing discipline, not a one-time project. Runbook Development & Knowledge Management --- Own the platform's operational runbooks --- ensuring every known failure mode, recovery procedure, and escalation path is documented, version-controlled, regularly tested, and immediately accessible during an incident. The runbook library must be current, complete, and something an engineer can actually use under pressure at 2am. Capacity Planning --- Work with the Platform Manager and Cloud Engineers to continuously monitor platform resource utilisation and anticipate capacity requirements before they become constraints. Ensure the platform can absorb growth --- in users, traffic, and data volume --- ahead of demand, not in response to it. Team Leadership --- Lead, mentor, and develop the Site Reliability Engineers on the team --- setting clear expectations, reviewing their work, and creating the conditions for them to grow into increasingly independent, senior practitioners. Build a high-performing SRE team that the rest of the engineering organisation trusts and relies on. What We're Looking For Must Have 5+ years of SRE, DevOps, or platform engineering experience --- with at least 2 years in a lead or senior individual contributor role owning a reliability function in production. Deep hands-on experience with observability tooling --- centralised logging, metrics (Prometheus, Datadog, or equivalent), distributed tracing (Jaeger, Open Telemetry, or equivalent), and alerting frameworks. You have built and operated an observability stack, not just used one. Proven experience designing and running incident management processes --- including on-call frameworks, severity classification, escalation paths, blameless post-mortems, and action item tracking. Strong understanding of SRE principles --- SLOs, SLAs, error budgets, toil reduction, and reliability engineering as a discipline distinct from DevOps. Hands-on experience with cloud infrastructure on AWS and/or Azure --- including compute, managed services, networking, and the observability capabilities of both platforms. Experience with DORA metrics --- deployment frequency, lead time, change failure rate, MTTR --- and using engineering performance data to drive improvement. Strong scripting and automation skills --- Bash, Python, or equivalent --- for building the tooling and automation that makes the platform easier to operate at scale. Nice to Have Experience designing and running chaos engineering programmes --- Chaos Monkey, Gremlin, or equivalent --- and using failure injection to validate platform resilience. Experience building and operating a 24/7 NSOC or SOC function in a regulated financial services environment. Familiarity with container and Kubernetes operational patterns --- pod health, cluster observability, resource quotas, and Kubernetes-native alerting. Experience with capacity planning and FinOps --- using utilisation data to inform infrastructure scaling and cost decisions. AWS and/or Microsoft Azure certification --- SysOps, DevOps, or Solutions Architect tracks. Experience working in fintech, banking, or a regulated environment --- with an understanding of the reliability and availability expectations that financial services demand.

Similar jobs

Similar jobs

株式

fixed-term project engagement|Lead DevOps / SRE

株式会社天地人

🌍WorldwideJul 15, 2026, 8:04 AM UTC
Plata Card logo

QA Engineer - Fullstack(Mobile) Junior+

Plata Card

🌍Worldwide11 hours ago
Plata Card logo

QA Engineer - Backend (Golang) Junior+ [KYC]

Plata Card

🌍Worldwide11 hours ago
MOLO17 srl logo

Product Engineer (Mid), Java / Kotlin

MOLO17 srl

🌍Worldwide12 hours ago
MOLO17 srl logo

Product Engineer (Junior), Java / Kotlin

MOLO17 srl

🌍Worldwide12 hours ago
Careerflow logo

AI/ML Software Engineer-Task Creator (RL Environments) (Contract)

Careerflow

🌍Worldwide12 hours ago