Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
ITDS logo

Site Reliability Engineer

ITDS
Posted 2 days ago
🇵🇹Portugal🏠Remote📁Engineering & Development
Is this job info correct?

Empower Cloud Resilience — Drive Innovation and Reliability at Scale!


Portugal-based opportunity with a fully remote work arrangement (up to 5 days per week).


As a Site Reliability Engineer (AWS), you will own the reliability, scalability, and operational health of our client's AWS-based platform — a leading player in the cloud and infrastructure domain. This is a hands-on engineering role with a systems mindset: you'll define and defend SLOs, drive down toil through automation, lead incident response, and build the tooling and infrastructure-as-code that lets development teams ship safely and fast.


Your main responsibilities:

  • Define, measure, and enforce SLOs, SLIs, and error budgets, and use them to drive engineering priorities and release decisions.
  • Own the incident lifecycle — detection, response, mitigation, and blameless post-mortems — reducing mean time to recovery through better tooling and runbooks.
  • Design and maintain highly available, fault-tolerant AWS architectures across compute, networking, storage, and data layers.
  • Build and maintain infrastructure-as-code and CI/CD pipelines that make deployments repeatable, auditable, and low-risk.
  • Systematically identify and eliminate toil through automation, replacing manual operational work with self-healing systems.
  • Own observability — metrics, logging, tracing, alerting — so problems surface before they reach customers.
  • Partner with development teams on capacity planning, performance tuning, cost optimization, and production-readiness reviews.
  • Improve the security and compliance posture of the cloud environment in collaboration with security teams.
  • Participate in and improve the on-call rotation, tuning alerts to reduce noise and fatigue.


You're ideal for this role if you have:

  • Strong production experience operating systems at scale on AWS, including core services (EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch).
  • Solid grounding in SRE principles — SLOs, error budgets, toil reduction, blameless post-mortems.
  • Hands-on expertise with infrastructure-as-code (Terraform preferred; CloudFormation or CDK acceptable).
  • Proficiency with container orchestration (Kubernetes / EKS) and CI/CD tooling.
  • Strong scripting and automation skills in at least one language (Python, Go, or Bash).
  • Deep experience with observability stacks (Prometheus, Grafana, Datadog, ELK, or equivalents).
  • Demonstrated ownership of incident response and on-call in a production environment.
  • Sound understanding of networking, Linux internals, and distributed-systems failure modes.

Nice to have (not required):

  • AWS certifications (Solutions Architect, DevOps Engineer, or SysOps).
  • Experience with multi-account / multi-region AWS setups and cost governance.
  • Chaos engineering or resilience-testing experience.
  • Service mesh, GitOps (ArgoCD / Flux), or policy-as-code (OPA).
  • Regulated-industry background (financial services, healthcare).


Language: Fluent English


Eligibility:

  • Only candidates with a legal right to work in Portugal or the wider European Union will be considered.
  • Only candidates based in Portugal will be considered


#MAKEYourCareerBETTER


Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.

Similar jobs

Similar jobs

Inetum logo

DevOps / Site Reliability Engineer (SRE)

Inetum

🇵🇹PortugalYesterday
ITDS logo

Senior AWS Site Reliability Engineer - Remote

ITDS

🇵🇹Portugal2 days ago
IgniteTech logo

AI-DNA Senior Infrastructure & Reliability Engineer

IgniteTech

🌍Poland, Portugal, Ukraine1 weeks ago
Paddle logo

Site Reliability Engineer

Paddle

🌍Ireland, Portugal, United KingdomJun 23, 2026, 7:22 PM UTC
GoCardless logo

Site Reliability Engineer

GoCardless

🌍Portugal, United KingdomJun 17, 2026, 8:10 AM UTC
Intapp logo

Site Reliability Engineer

Intapp

🌍Portugal, United KingdomMay 30, 2026, 6:33 PM UTC