RITS Professional Services logo

Service Manager/Site Reliability Engineer

Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 28, 2026
Is this job info correct?

Service Manager/Site Reliability Engineer

We are looking for Service Manager/Site Reliability Engineer.

Rate: 160-180 PLN/h net + VAT (B2B)

100% remote

Project Description
Join a global technology organization focused on ensuring the reliability, stability, and operational excellence of large-scale production systems. The role sits at the intersection of Incident Operations, Site Reliability Engineering (SRE), and technical stakeholder communication, supporting real-time incident management, impact assessment, and operational improvements in a fast-paced, highly available environment. You will work closely with engineering and operational teams to maintain service reliability, improve incident processes, and drive automation initiatives

Responsibilities:

  • Monitor, triage, and coordinate responses to production incidents and operational alerts.
  • Act as a central coordination point between engineering teams and key stakeholders during incidents.
  • Assess incident impact, determine severity, and coordinate communications according to SLA commitments.
  • Manage incident lifecycles from detection through resolution and post-incident activities.
  • Maintain external-facing incident communications and status updates.
  • Support incident reporting, root cause analysis (RCA), and operational reviews.
  • Contribute to process improvements, automation initiatives, and operational tooling enhancements.
  • Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.
  • Participate in reliability-focused development activities and support operational excellence initiatives.

We are looking for:

  • 5+ years of experience in Incident Operations, Site Reliability Engineering (SRE), Technical Operations, or a similar role.
  • Experience working in on-call environments with SLA-driven responsibilities.
  • Proven ability to operate effectively during high-pressure, real-time incident scenarios.
  • Strong understanding of distributed systems and production environments.
  • Experience with monitoring and alerting tools such as Datadog, Chronosphere, or similar.
  • Experience with incident management platforms such as PagerDuty, Rootly, or comparable tools.
  • Familiarity with APIs, system integrations, observability tools, and monitoring dashboards.
  • Hands-on programming experience with Python and/or Kotlin.
  • Understanding of the Software Development Lifecycle (SDLC) and production reliability principles.
  • Excellent written and verbal communication skills.
  • Strong operational judgment and ability to make decisions under uncertainty.
  • Ability to manage multiple priorities simultaneously in a fast-paced environment.
  • Highly organized with strong ownership and attention to detail.
  • Strong collaboration skills with cross-functional engineering and business teams.
  • Proactive mindset focused on continuous improvement, automation, and scalability.

This role is not perfectly suited for you, but you have a friend who would fit? Recommend your friend and get up to 5000 zł!
Referral Program: Talent from your network


Don't hesitate and apply now!

Similar jobs

Apply for this job