Xideral logo

Senior AWS Site Reliability Engineers

Xideral
Posted 12 hours ago
MexicoHybridEngineering & Development
Is this job info correct?

Seeking Experienced Senior AWS Site Reliability Engineers for Exciting Projects – Remote in Mexico

We are looking for skilled Senior Site Reliability Engineers with a minimum of 5 years of experience to join a dynamic team within a leading organization. This role involves supporting and improving cloud operations for microservice-based platforms, with a focus on production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments.

Key Responsibilities:

  • Own and improve the reliability of cloud-based services and supporting infrastructure.
  • Participate in on-call rotations and support production systems outside normal business hours.
  • Lead incident response activities, including triage, escalation, mitigation, and service restoration.
  • Drive blameless postmortems and ensure corrective actions are tracked to closure.
  • Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis.
  • Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools.
  • Support and improve cloud and container platforms across AWS and Azure.
  • Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems.
  • Build automation to reduce manual effort and improve operational efficiency.
  • Configure and improve monitoring, alerting, logging, diagnostics, and observability.

Technical Skills Required:

With over 5 years of experience as a Senior Site Reliability Engineer, you must be proficient in the following technical skills:

  • Strong hands-on experience with AWS and Azure cloud platforms.
  • Strong experience with Terraform for Infrastructure as Code (IaC).
  • Experience with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools.
  • Strong hands-on experience with Docker and Kubernetes.
  • Experience designing, maintaining, and troubleshooting complex CI/CD pipelines.
  • Strong production support experience, including incident management, Root Cause Analysis (RCA), postmortems, and runbook creation.
  • Strong observability experience, including monitoring, alerting, logging, diagnostics, and performance analysis.
  • Good understanding of cloud networking, security, access controls, and InfoSec practices.
  • Experience with version control, branching, merging, pull requests, and conflict resolution.
  • Understanding of cloud cost optimization and resource utilization.

Good-to-Have Skills:

  • Experience with microservice-based platforms.
  • Experience with Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar tools.
  • Scripting or programming experience using Python, Bash, Go, or Java.
  • Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering.
  • Experience with disaster recovery testing and production readiness reviews.
  • Prior experience mentoring junior engineers or leading technical troubleshooting.

Qualifications:

  • Bachelor’s degree or higher.
  • Fluent in English (Advanced).
  • Excellent communication, empathy, commitment, leadership, teamwork, and a proactive attitude.

Location & Schedule:

  • Remote work from Mexico.
  • Preferred hybrid model in Guadalajara, Jalisco, with expected onsite attendance 2 days per week.
  • Work hours Monday to Friday, 09:00 – 18:00.
  • Advanced English skills are mandatory, and only residents of Mexico.

Benefits:

  • Attractive Salary + Premium Benefits
  • Performance bonuses, grocery coupons, and savings are found.
  • Aguinaldo, premium vacations, and vacations paid
  • SGMM Medical insurance, family, and Life insurance.

Candidates must include their compensation expectations in their applications and resumes in English.

Interested? Apply now through this link:

Similar jobs