Observability & SRE Engineer
Mode: 1 year Fixed term contract
Remote but candidate based in Germany are preferred
German Profecincy minimum of C1 is mandatory
Primary Skills:
Prometheus, Grafana, Loki, OpenTelemetry SDKs, Jaeger tracing, SIEM log export.
Security Clearance & Vetting Level: Public Sector Clearance + NdK (Nachweis der Kundigkeit)
Position Overview
The Observability & SRE Engineer builds and operates central telemetry stacks to provide visibility across distributed cloud infrastructure. This role implements metric collection, log aggregation, distributed tracing, and standalone alerting tailored for Air-Gap operations.
Technical Qualifications & Skills
• Must Have:
o Deep expertise in Prometheus (PromQL, ServiceMonitor, Federation, Remote Write) and Grafana.
o Hands-on experience with Loki log aggregation and AlertManager routing.
o Proficiency with OpenTelemetry (Collectors, SDKs, OTLP) and Jaeger distributed tracing.
o Knowledge of SIEM integrations and security event logging.
o Scripting skills in Go, Python, Shell, and YAML.
Site Reliability Engineer (m/f/d)
gridscale GmbH
Site Reliability Engineer II - AI & Infrastructure (f/m/d)
Focused
Senior Platform Operations Engineer / Site Reliability Engineer (m/w/d)
Netlution GmbH
Site Reliability Engineer (f/m/d) – Observability & Internal Tools
Bertelsmann
Site Reliability Engineer (f/m/d) – Observability & Internal Tools
RTL Deutschland
Senior Site Reliability Engineer (Performance and Scalability)
Digital Zone