KL
Site Reliability Engineer
Kiberon LabsFull-time
Kiberon Labs is a specialized AI and DevSecOps consultancy. We help clients design, secure, migrate, and operate systems that need to perform reliably in production.
Our work includes solution architecture, process migration, security hardening, Kubernetes operations, and forward-deployed engineering. We work directly with client teams to solve practical infrastructure problems and leave them with systems they can confidently own.
We are a small technical team. That means less distance between an idea, a decision, and the system running in production.
The role
We are hiring one full-time Site Reliability Engineer to design, operate, and improve reliable full-stack systems running at scale on Kubernetes.
You will work directly with the founder, software engineers, and external client teams. Your job is to make production systems more observable, resilient, secure, and easier to operate.
This is a hands-on role. You will investigate failures, automate repetitive operations, improve infrastructure, and help teams adopt practical SRE methods.
We are not assigning a rigid seniority label to this position. We care about your ability to operate complex production systems and take responsibility for technical outcomes.
What you will do
This is a remote position for candidates located within two time zones of GMT+1.
Our fixed core collaboration hours are 10:00 to 14:00 GMT+1. Outside that window, you have flexibility to structure your working day around the work and your responsibilities.
You will have direct access to the founder and meaningful input into technical decisions. We expect you to challenge weak assumptions, propose better approaches, and take ownership of the systems you help build.
What we offer
Interested?
If you enjoy owning difficult production problems and building systems that remain reliable under real operating conditions, apply through LinkedIn.
- Remote
- Within two time zones of GMT+1
Kiberon Labs is a specialized AI and DevSecOps consultancy. We help clients design, secure, migrate, and operate systems that need to perform reliably in production.
Our work includes solution architecture, process migration, security hardening, Kubernetes operations, and forward-deployed engineering. We work directly with client teams to solve practical infrastructure problems and leave them with systems they can confidently own.
We are a small technical team. That means less distance between an idea, a decision, and the system running in production.
The role
We are hiring one full-time Site Reliability Engineer to design, operate, and improve reliable full-stack systems running at scale on Kubernetes.
You will work directly with the founder, software engineers, and external client teams. Your job is to make production systems more observable, resilient, secure, and easier to operate.
This is a hands-on role. You will investigate failures, automate repetitive operations, improve infrastructure, and help teams adopt practical SRE methods.
We are not assigning a rigid seniority label to this position. We care about your ability to operate complex production systems and take responsibility for technical outcomes.
What you will do
- Design and operate scalable Kubernetes infrastructure.
- Improve the reliability and availability of client applications and services.
- Build and maintain monitoring, logging, tracing, and alerting systems.
- Investigate production incidents across application and infrastructure layers.
- Improve incident response procedures and contribute to post-incident reviews.
- Automate operational tasks using code, scripts, and infrastructure-as-code tooling.
- Identify reliability risks before they become production failures.
- Improve deployment pipelines, configuration management, and release safety.
- Diagnose problems across networking, storage, compute, security, and application runtimes.
- Help engineering teams define practical service objectives and reliability measures.
- Document technical decisions and transfer operational knowledge to client teams.
- Contribute to internal technical systems and product development when opportunities arise.
- Strong hands-on experience operating Kubernetes in production.
- Experience supporting large-scale, full-stack production systems.
- Strong AWS infrastructure and operations knowledge.
- Experience with observability, monitoring, alerting, and production diagnostics.
- Experience responding to incidents and troubleshooting distributed systems.
- Strong Linux system administration and performance investigation skills.
- Working knowledge of networking, storage, security, and cloud architecture.
- Software development or scripting skills for automation and internal tooling.
- Experience with infrastructure-as-code and configuration management.
- The ability to explain technical risks and decisions clearly.
- The ability to collaborate with both technical and non-technical stakeholders.
- A practical approach to reliability, security, and operational complexity.
- TypeScript or Go.
- Azure infrastructure.
- CI/CD pipeline design and operation.
- DevSecOps and secure configuration management.
- Security hardening for cloud-native systems.
- Client-facing consulting or embedded engineering.
- Cost optimization and capacity planning.
- Designing service-level indicators and objectives.
This is a remote position for candidates located within two time zones of GMT+1.
Our fixed core collaboration hours are 10:00 to 14:00 GMT+1. Outside that window, you have flexibility to structure your working day around the work and your responsibilities.
You will have direct access to the founder and meaningful input into technical decisions. We expect you to challenge weak assumptions, propose better approaches, and take ownership of the systems you help build.
What we offer
- Direct founder access and short decision paths.
- Flexible working hours outside the core collaboration window.
- A monetary learning budget for courses, books, certifications, and technical resources.
- Dedicated working time for learning, experimentation, and testing new ideas.
- Opportunities to take technical ownership of internal systems and products.
- Work focused on modern infrastructure without unnecessary legacy bloat.
- The opportunity to help shape engineering practices inside a growing technical company.
- Potential access to non-voting profit-sharing shares through the employee pool.
Interested?
If you enjoy owning difficult production problems and building systems that remain reliable under real operating conditions, apply through LinkedIn.