Python Developer (SRE)
- Salary
- PLN 160–PLN 180/hrPLN per hour
- Hiring from
- Probably Worldwide
- Work type
- Remote
- Posted
- Sep 28, 2026
Python Developer (SRE)
We are looking for Python Developer (SRE).
Rate: 160-180 PLN/h net + VAT (B2B)
100% remote
Project Description
Join a global technology organization focused on ensuring the reliability, stability, and operational excellence of large-scale production systems. The role sits at the intersection of Incident Operations, Site Reliability Engineering (SRE), and technical stakeholder communication, supporting real-time incident management, impact assessment, and operational improvements in a fast-paced, highly available environment. You will work closely with engineering and operational teams to maintain service reliability, improve incident processes, and drive automation initiatives
Responsibilities:
- Monitor, triage, and coordinate responses to production incidents and operational alerts.
- Act as a central coordination point between engineering teams and key stakeholders during incidents.
- Assess incident impact, determine severity, and coordinate communications according to SLA commitments.
- Manage incident lifecycles from detection through resolution and post-incident activities.
- Maintain external-facing incident communications and status updates.
- Support incident reporting, root cause analysis (RCA), and operational reviews.
- Contribute to process improvements, automation initiatives, and operational tooling enhancements.
- Collaborate with engineering teams to improve observability, monitoring, and incident response capabilities.
- Participate in reliability-focused development activities and support operational excellence initiatives.
We are looking for:
• 5+ years of hands-on experience in Software Engineering, Site Reliability Engineering (SRE), Production Engineering, Incident Operations, or related technical roles.
• Strong software development background with recent, demonstrable experience building and maintaining production-grade applications and automation in Python.
• Experience working in on-call environments with SLA/SLO-driven operational responsibilities.
• Proven ability to operate effectively during high-severity, real-time production incidents.
• Solid understanding of distributed systems, cloud-native architectures, and large-scale production environments.
• Experience troubleshooting complex application, infrastructure, and service reliability issues.
• Advanced proficiency in Python development, including building automation, tooling, integrations, and operational services.
• Experience with software engineering best practices including testing, code reviews, CI/CD, and version control.
• Strong understanding of the Software Development Lifecycle (SDLC) and production reliability engineering principles.
• Familiarity with Kotlin is a plus.
• Experience with Slack automation and operational workflows is desirable. Incident Response & Reliability.
• Hands-on experience coordinating, managing, and resolving production incidents.
• Experience assessing customer impact, driving remediation efforts, and leading technical investigations.
• Ability to create and execute operational runbooks and automate repetitive operational tasks.
• Strong understanding of observability, monitoring, alerting, and incident response processes.
• Experience performing root cause analysis and driving continuous reliability improvements.
• Monitoring and observability platforms (e.g., Datadog, Chronosphere) o Incident management platforms (e.g., PagerDuty, Rootly)
• APIs and service integrations
• Production debugging and root cause analysis
• Reliability engineering concepts including SLI/SLOs, error budgets, toil reduction, and automated remediation
This role is not perfectly suited for you, but you have a friend who would fit? Recommend your friend and get up to 5000 zł!
Referral Program: Talent from your network
Don't hesitate and apply now!