Strategic Staffing Solutions logo

Site Reliability Engineer

Hiring from
Argentina
Work type
Remote
Posted
Is this job info correct?
Show job description

We are looking for an experienced Site Reliability Engineer (SRE) to ensure the availability, performance, and resilience of critical systems in a global enterprise environment.


This is a hands-on SRE role focused on proactive incident prevention, production monitoring, automation, and continuous improvement of service reliability. The ideal candidate will work closely with product line agile teams, infrastructure and security teams, and SRE leadership to reduce toil, improve observability, and drive SRE maturity across the organization.


🎯 Key Responsibilities

  • Prevent incidents proactively by baselining expected service performance, applying lessons learned, and using data analytics to identify problem areas and operational gaps.
  • Lead product line agile teams in troubleshooting and resolving system problems, including analysis of application and critical system performance.
  • Serve as a technical resource during critical and major incidents across multiple technologies.
  • Facilitate SRE technical assessments, identify gaps, and provide recommendations to product teams on their SRE maturity journey based on the organization's SRE framework.
  • Improve logging and create automated resolutions based on triggers to avoid future issues.
  • Develop automation scripts for repetitive tasks to eliminate toil and operations support activities.
  • Oversee production environments by monitoring availability and maintaining a holistic view of system health.
  • Measure and optimize system performance, continuously seeking innovation and improvement to meet customer needs.
  • Collaborate with peers, company leadership, subject matter experts, and users to promote end-to-end DevOps / SRE best practices.
  • Partner with the SRE Community of Practice to define SRE capabilities and best practices, and integrate the capability framework throughout the organization.


🛠️ Required Qualifications

  • Hands-on experience as an IT professional with full stack infrastructure knowledge and experience troubleshooting incidents and production issues.
  • Working knowledge of network administration & security (Cisco/Juniper), identity & access management (Active Directory, Azure AD, SAML, OpenID Federation, certificates, and keys), cybersecurity, on-prem & cloud architecture, Windows & Linux OS, performance monitoring, application & database troubleshooting, change management, and API integration.
  • Experience with automation (Ansible, PowerShell, KQL, or Shell scripting).
  • Experience supporting critical and major incident response.
  • Strong written and verbal communication and facilitation skills, with the ability to convey business and technical information to a diverse audience.
  • Strong analytical and problem-solving skills, with persistence in resolving difficult issues.
  • Advanced English (mandatory).


⭐ Nice to Have

  • Experience applying SRE practices such as SLIs/SLOs, error budgets, and postmortems.
  • Experience with observability and monitoring platforms.
  • Experience conducting SRE maturity assessments.
  • Experience supporting production environments in 24/7 operations.
  • Experience working with global enterprise organizations.
  • Experience in the Oil & Gas industry.
  • Relevant cloud or SRE certifications.


💎 What We Offer

  • Full time employment.
  • 100% remote.
  • Competitive salary in ARS.
  • Opportunity to work on global enterprise projects with an international team.

Similar jobs

Apply on LinkedIn