PA

Senior Infrastructure Reliability Specialist

Hiring from
United Arab Emirates
Work type
Remote
Posted
Is this job info correct?

507,722 remote jobs, straight from company career pages

100% free · New jobs every hour

Show job description
We're Hiring: Senior Infrastructure Reliability Specialist

Location: United Arab Emirates (Remote)

Employment Type: Full-Time

Experience Level: Senior

Work Arrangement: Fully Remote

About Us

We are a globally focused organization committed to delivering reliable, secure, scalable, and resilient technology services across diverse markets. Our teams combine Infrastructure Engineering, Cloud Operations, Site Reliability Engineering, Cybersecurity, Network Engineering, Architecture, and Technology Operations to maintain dependable digital platforms and business-critical systems.

We emphasize proactive monitoring, automation, measurable service reliability, operational excellence, and continuous improvement to ensure our technology infrastructure supports business growth and uninterrupted service delivery.

The Role

We are seeking an experienced Senior Infrastructure Reliability Specialist to improve the availability, performance, resilience, and operational reliability of enterprise infrastructure across cloud, on-premises, network, storage, compute, and hybrid environments.

The ideal candidate will combine infrastructure engineering expertise with reliability engineering, observability, automation, incident management, capacity planning, and problem management to minimize service disruptions, eliminate recurring failures, and establish measurable reliability standards across critical technology services.

Key Responsibilities
  • Develop and implement infrastructure reliability strategies, standards, operating procedures, and engineering practices.
  • Establish reliability objectives, service-level indicators (SLIs), service-level objectives (SLOs), and service-level agreements (SLAs) for infrastructure services.
  • Define and monitor infrastructure availability, performance, resilience, capacity, and recovery targets.
  • Assess infrastructure architectures and identify reliability risks, single points of failure, and operational weaknesses.
  • Design and implement resilient infrastructure solutions across cloud, on-premises, and hybrid environments.
  • Monitor the health, availability, performance, and resource utilization of servers, virtual machines, containers, storage platforms, networks, and cloud services.
  • Implement and maintain infrastructure observability, monitoring, logging, alerting, and diagnostic capabilities.
  • Develop actionable alerts that identify infrastructure degradation before it affects business services.
  • Reduce alert noise by improving thresholds, correlation, prioritization, and automated event enrichment.
  • Analyze infrastructure metrics, logs, traces, events, and performance data to identify emerging reliability issues.
  • Lead root-cause analysis for infrastructure incidents, recurring failures, service degradation, and unexpected outages.
  • Develop corrective and preventive actions to eliminate recurring incidents and improve system stability.
  • Maintain problem records, known-error documentation, technical investigations, and remediation plans.
  • Coordinate major incident response involving infrastructure failures, performance degradation, network interruptions, storage issues, or cloud-service disruptions.
  • Support incident command, technical troubleshooting, stakeholder communications, recovery coordination, and post-incident reviews.
  • Establish clear escalation procedures and technical ownership for infrastructure reliability issues.
  • Develop and maintain infrastructure runbooks, operational playbooks, recovery procedures, and troubleshooting guides.
  • Implement automation for routine infrastructure operations, health checks, remediation, deployment, and recovery activities.
  • Develop scripts and automation workflows using appropriate tools and languages such as Python, PowerShell, Bash, or configuration-management frameworks.
  • Apply Infrastructure as Code practices to improve consistency, repeatability, and reliability of infrastructure provisioning.
  • Support configuration management, version control, infrastructure baselines, and controlled changes.
  • Improve infrastructure deployment and change processes to reduce failure rates and minimize service disruption.
  • Conduct risk assessments for infrastructure changes, maintenance activities, platform upgrades, and technology migrations.
  • Coordinate patching, upgrades, hardware refreshes, firmware updates, and lifecycle management with relevant engineering teams.
  • Monitor infrastructure configuration drift and ensure systems remain aligned with approved technical standards.
  • Implement high-availability architectures, redundancy, failover mechanisms, and fault-tolerant designs.
  • Validate load balancing, clustering, replication, backup, and recovery mechanisms.
  • Develop and maintain infrastructure disaster-recovery plans, technical recovery procedures, and service-restoration priorities.
  • Coordinate disaster-recovery exercises, failover testing, recovery validation, and resilience assessments.
  • Verify recovery time objectives (RTOs) and recovery point objectives (RPOs) for critical infrastructure services.
  • Assess backup integrity, restoration performance, replication health, and recovery readiness.
  • Support business continuity planning for infrastructure-dependent services.
  • Conduct capacity planning and forecasting for compute, storage, memory, network bandwidth, cloud resources, and platform demand.
  • Identify resource constraints, performance bottlenecks, saturation risks, and emerging capacity issues.
  • Recommend infrastructure scaling, resource optimization, workload distribution, and performance improvements.
  • Monitor infrastructure costs and identify opportunities to improve utilization and operational efficiency.
  • Evaluate infrastructure reliability risks associated with cloud adoption, migrations, modernization, and platform consolidation.
  • Collaborate with Cloud Engineering, Network Engineering, Security, DevOps, Application Support, and Architecture teams to resolve cross-platform reliability issues.
  • Establish technical standards for infrastructure health checks, resilience testing, monitoring coverage, and operational readiness.
Key Performance Indicators
  • Infrastructure availability
  • Critical service uptime
  • Infrastructure reliability SLO attainment
  • SLA compliance
  • Infrastructure incident frequency
  • Major incident frequency
  • Repeat incident rate
  • Mean time to detect (MTTD)
  • Mean time to acknowledge (MTTA)
  • Mean time to recover (MTTR)
  • Mean time between failures (MTBF)
  • Infrastructure change failure rate
  • Incident recurrence reduction
  • Root-cause analysis completion
  • Corrective-action closure rate
  • Monitoring coverage
  • Alert accuracy
  • False-positive alert rate
  • Alert noise reduction
  • Infrastructure health-check compliance
  • Capacity utilization efficiency
  • Capacity forecast accuracy
  • Infrastructure performance benchmark compliance
  • Resource saturation incident rate
  • Cloud and infrastructure cost efficiency
  • Automation coverage
  • Automated remediation success rate
  • Infrastructure provisioning consistency
  • Configuration drift resolution
  • Patch and upgrade compliance
  • High-availability test success rate
  • Failover test success rate
  • Disaster-recovery test completion
  • Recovery time objective compliance
  • Recovery point objective compliance
  • Backup success rate
  • Restore-test success rate
  • Replication health
  • Infrastructure vulnerability remediation timeliness
  • Production change success rate
  • Infrastructure documentation completeness
  • Dependency mapping accuracy
  • Operational runbook coverage
  • Reliability improvement initiative completion
  • Infrastructure technical-debt reduction
  • Audit and control compliance
  • Stakeholder satisfaction
  • Engineering knowledge-sharing and training completion
Ideal Candidate

The successful candidate should have strong experience in infrastructure reliability engineering, site reliability engineering, infrastructure operations, cloud engineering, systems engineering, or enterprise technology operations, preferably within a complex, business-critical, or highly available technology environment.

The candidate should demonstrate:

  • Strong knowledge of infrastructure architecture, reliability engineering, and operational resilience.
  • Proven experience managing reliability across cloud, on-premises, or hybrid infrastructure environments.
  • Strong understanding of compute, storage, virtualization, networking, operating systems, and cloud platforms.
  • Experience with AWS, Microsoft Azure, Google Cloud, or comparable enterprise infrastructure environments.
  • Strong understanding of infrastructure monitoring, observability, logging, metrics, and alerting.
  • Experience with monitoring platforms such as Prometheus, Grafana, Datadog, Dynatrace, New Relic, or equivalent tools.
  • Strong incident management, problem management, and root-cause analysis capabilities.
  • Experience handling major infrastructure incidents and coordinating technical recovery activities.
  • Strong knowledge of high availability, redundancy, failover, fault tolerance, and disaster recovery.

Similar jobs

Apply on LinkedIn