Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
CXM logo

Site Reliability Engineer

CXM
Posted 1 weeks ago
🌍Latin America, North America🏠Remote📁Engineering & Development
Is this job info correct?

Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE) , you will own the day-to-day reliability of our .NET/C# services running on Windows , starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services. This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience. Position Details Team Platform & Production Reliability Location Remote (Americas, LatAm preferred) Working Hours Americas time zones (UTC-3 to UTC-8) On-call Rotation aligned with the London trading day Employment Type Full-time, Permanent Experience Level Mid-Level (3–5 years) Technology Stack .NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform About the Role Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient. You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault isolation, and incident response, while driving continuous improvements in platform reliability and operational excellence. What You'll Do Participate in the on-call rotation for production trading systems and lead incident response during service disruptions. Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues. Build and maintain Grafana dashboards , Prometheus alerts , and operational health views across applications, infrastructure, and databases. Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact. Define, implement, and monitor Service Level Indicators (SLIs) , Service Level Objectives (SLOs) , and error budgets. Troubleshoot issues across: .NET/C# applications Windows Server Aurora PostgreSQL databases AWS infrastructure CI/CD pipelines and deployments Improve deployment safety, release automation, and rollback strategies. Partner with developers to improve application operability, resilience, and fault isolation. Automate operational tasks through scripting and infrastructure automation. Create and maintain runbooks, operational documentation, and incident response procedures. Continuously improve monitoring, alert quality, automation, and platform reliability. Required Technical Skills .NET & Windows Strong experience debugging and supporting .NET/C# applications in production. Hands-on experience with Windows Server environments. Scripting & Automation Strong PowerShell scripting skills. Experience with Python or Bash . Observability Experience with Grafana , Prometheus , and Loki (or equivalent monitoring and observability tools). Solid understanding of metrics, logging, tracing, and alerting best practices. CI/CD & DevOps Experience with modern CI/CD pipelines. Knowledge of deployment strategies, release automation, and rollback mechanisms. Cloud & Infrastructure Experience working with AWS . Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools. Databases Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms. Reliability Engineering Practical experience with: SLIs & SLOs Error Budgets Incident Response Root Cause Analysis (RCA) Alert Design Production Operations Preferred Qualifications Experience supporting high-availability or low-latency financial or trading systems. Familiarity with MetaTrader environments or financial technology platforms. Experience with distributed systems and microservices. Knowledge of OpenTelemetry or similar observability frameworks. Exposure to Docker, Kubernetes, or containerized environments. Why Join Us? Work on mission-critical trading infrastructure that directly impacts customers. Solve challenging reliability and scalability problems in a real-time environment. Build world-class observability, automation, and deployment practices. Collaborate with experienced engineers in a modern engineering culture. Influence reliability strategy and engineering best practices across the platform. If you're passionate about production engineering, automation, and building reliable systems at scale, we'd love to hear from you.

Similar jobs

Similar jobs

Gradle Technologies logo

Staff Site Reliability Engineer

Gradle Technologies

🌍Europe, North America20 hours ago
Gradle Technologies logo

Senior Site Reliability Engineer

Gradle Technologies

🌍Europe, North America20 hours ago
BairesDev logo

Senior Site Reliability Engineer (SRE) - Remote Work

BairesDev

🌍Latin America3 days ago
Sezzle logo

Senior Site Reliability Engineer

Sezzle

🌍India, Latin America, Turkey3 days ago
Truelogic logo

Senior Site Reliability Operations Engineer - Finance

Truelogic

🌍Latin America3 days ago
Sezzle logo

Principal Site Reliability Engineer

Sezzle

🌍Latin America1 weeks ago