Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Telekom Growthhub logo

Site Reliability Engineer (SRE) (m/f/d)

Telekom Growthhub
Posted 3 weeks ago
🇩🇪Germany🏢Hybrid📁Engineering & Development
Is this job info correct?

About T-Systems International GmbH At T-Systems, we offer business customers the right system solutions for their digital business. With our portfolio we ensure that digital transformation reduces complexity, saves costs and makes day-to-day work easier. We focus on the areas connectivity, digital, cloud & infrastructure as well as security - Let's power higher performance! About the Position As a Site Reliability Engineer , you are at the forefront of ensuring the stability, performance, and scalability of our mission-critical systems. You play a key role in the design, implementation, and maintenance of a highly reliable and efficient infrastructure. Leveraging your expertise in modern monitoring, observability, and cloud-native technologies, you design, build, and operate a high-performance monitoring stack based on Prometheus, Grafana, Loki, and Alertmanager for distributed Fog and Edge infrastructures. You develop specialized dashboards for monitoring hardware health, resource utilization, and network connectivity, thereby establishing the foundation for transparent and reliable operations. You design and implement alerting strategies for autonomous operating models as well as disconnected and air-gapped environments, ensuring that critical events are reliably detected and processed even under limited network connectivity. In addition, you are responsible for implementing local data retention concepts, log rotation mechanisms, and efficient data management strategies tailored to resource-constrained environments. Another key focus of your role is the continuous performance optimization of the monitoring stack while taking into account the limited CPU, memory, and storage resources available on Fog nodes. You integrate Kubernetes monitoring solutions, analyze operational and performance metrics, and proactively identify optimization opportunities to improve the platform's stability, availability, and overall efficiency. By applying Infrastructure as Code (IaC) principles and YAML-based configurations, you automate the deployment and management of monitoring components and make a significant contribution to the continuous evolution of a robust, scalable, and highly available observability platform for modern Edge and Fog computing environments. Must-Have Skills Monitoring & Observability Strong experience with Prometheus , including: PromQL Federation Remote Write Local Retention Management Dashboarding & Visualization Advanced knowledge of Grafana , including: Dashboard development Alerting Provisioning as Code Log Management Experience with Loki , including: LogQL Retention Management Compaction Alerting & Incident Management Strong knowledge of Alertmanager , including: Standalone routing Alert inhibition Escalation and notification strategies Kubernetes & Cloud-Native Technologies Experience in monitoring and operating Kubernetes environments Solid understanding of container platforms and cloud-native architectures Edge & Fog Computing Experience working in resource-constrained environments Optimization of CPU, memory, and storage utilization Understanding of offline, air-gapped, and disconnected operating scenarios Additional Requirements Advanced proficiency in YAML Strong understanding of IT security principles and security awareness Strong analytical and structured approach to problem-solving Experience in Site Reliability Engineering (SRE) or infrastructure operations Excellent German language skills, both written and spoken (C1 level) Nice-to-Have Skills Experience with Infrastructure as Code (IaC) tools (e.g., Terraform, Ansible) Familiarity with GitOps methodologies and workflows Linux system administration Knowledge of OpenTelemetry and modern observability concepts Experience working with mission-critical or highly available infrastructures What You Can Expect You can expect an exciting role in the field of modern Edge, Fog, and cloud-native technologies . In this position, you will take ownership of the stability, availability, and observability of distributed infrastructures while actively contributing to the design and implementation of high-performance monitoring and observability solutions. Working in a technologically challenging environment, you will leverage state-of-the-art Kubernetes, monitoring, and logging technologies to help build robust and resilient platforms that ensure reliable operations—even under demanding conditions. What We Offer Flexibility and Work-Life Balance: Benefit from flexible working hours, mobile work options, and a hybrid work model. Attractive Compensation: A competitive salary package that recognizes your expertise and performance. Development Opportunities: Access to extensive training programs, courses, and career development initiatives. Modern Work Environment: Work in a modern setting equipped with the latest technologies and tools. Strong Company Culture: An open and collaborative work atmosphere where teamwork and mutual support are highly valued. Health and Well-being: Health promotion offers and initiatives that support your overall well-being. Additional Benefits: Various perks such as company pension scheme, employee discounts, and public transport ticket options. Your application may be processed by external service providers in Germany. Applications for positions outside Deutsche Telekom AG may also be processed by recruiters in other European countries. This also applies to the processing of applications at subsidiaries of Deutsche Telekom AG. Therefore, the following note applies to our civil servants: You are not required to include documents relevant to your personnel file, in particular civil service evaluations, with your application. However, you are free to provide these documents voluntarily. This position is also available part-time. People with disabilities will take priority in case of equal qualifications. For technical questions, please contact [email protected] https://www.xing.com/companies/deutschetelekomag Deutsche Telekom | LinkedIn

Similar jobs

Similar jobs

Exoscale logo

Site Reliability Engineer - Compute System & Network (f/m/d)

Exoscale

🌍Germany, SpainYesterday
Bertelsmann logo

Site Reliability Engineer (f/m/d) – Observability & Internal Tools

Bertelsmann

🇩🇪GermanyYesterday
TOPdesk logo

Senior Site Reliability Engineer (m/f/d)

TOPdesk

🇩🇪Germany3 days ago
Scalable GmbH logo

(Junior) Cloud Site Reliability Engineer (Network) (m/f/x)

Scalable GmbH

🇩🇪Germany3 days ago
Scalable GmbH logo

(Senior) Cloud Site Reliability Engineer (Scalability) (m/f/x)

Scalable GmbH

🇩🇪Germany3 days ago
Doctolib logo

Senior Site Reliability Engineer (x/f/m)

Doctolib

🌍France, Germany4 days ago