Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Systems Engineering Solutions Corporation logo

Reliability Engineer

Systems Engineering Solutions Corporation
Posted 3 weeks ago
🇺🇸United States🏠Remote📁Engineering & Development
Is this job info correct?

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer . The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission‑critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities. Location: This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA). System Reliability & Availability Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets Identify reliability risks and implement mitigation strategies across the system lifecycle Conduct capacity planning and performance modeling to ensure systems scale to meet demand Monitoring, Observability & Alerting Implement and manage monitoring, logging, and tracing solutions to provide full system observability Define actionable alerting thresholds that minimize noise and enable rapid incident detection Analyze trends and metrics to proactively identify potential reliability issues Incident Response & Problem Management Participate in on‑call rotations and lead incident response activities for production systems Coordinate troubleshooting efforts across development, infrastructure, and security teams Conduct post‑incident reviews (PIRs) and develop corrective and preventive action plans Track recurring issues and ensure root causes are resolved Automation & Engineering Excellence Automate operational tasks to reduce manual intervention and operational risk Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR) Promote “automation over toil” and standardize operational workflows Reliability‑Focused Engineering Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability Validate disaster recovery (DR) and business continuity plans; test failover mechanisms Support chaos engineering, fault injection testing, and resilience validation where appropriate Collaboration & Governance Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives Document system reliability standards, runbooks, and operational procedures Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls) Required Skills: · Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree. · Active Secret clearance at a minimum required to start · US citizenship required · Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services · Experience with containerized environments (Docker, Kubernetes) · Familiarity with CI/CD pipelines and deployment automation · SLOs and error budgets · Capacity modeling and performance testing · Strong understanding of: · Distributed systems and high‑availability architectures · Linux/Windows system administration · Networking fundamentals (DNS, TCP/IP, load balancing) · Hands-on experience with: · Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor) · Infrastructure as Code (Terraform, ARM, CloudFormation) · Scripting or programming languages (Python, Bash, Go, PowerShell, or similar) · Experience supporting incident management and on‑call operations Preferred Skills Experience with USAF Cloud One or Platform 1. Experience with Zero Trust Architecture Cloud certifications in AWS, Azure, Google, or Oracle clouds SES provides a competitive salary and the following benefits: Medical Dental Vision AD&D STD LTD Company paid Life Insurance 401k with employer contribution Paid Time Off Pet Insurance

Similar jobs

Similar jobs

Empower logo

Senior Data Reliability Engineer AWS

Empower

🇺🇸United States7 hours ago
KO

Senior Reliability Engineer (Remote)

Kohls

🇺🇸United States7 hours ago
DD

Site Reliability Engineer

Ddcdine

🇺🇸United States8 hours ago
itD Tech logo

Site Reliability Engineer (6266)

itD Tech

🇺🇸United States8 hours ago
Salas O'Brien logo

Reliability Engineer

Salas O'Brien

🇺🇸United States8 hours ago
I8

Mid-Senior Site Reliability Engineer – Kubernetes Platform

I8Is

🇺🇸United States8 hours ago