Intermedia Intelligent Communications logo
Hiring from
Portugal
Work type
Remote
Posted
Is this job info correct?
Show job description
*ALL CANDIDATES MUST BE LOCATED IN PORTUGAL*

About Intermedia

Are you looking for a company where YOUR VOICE is heard? Where can you MAKE A DIFFERENCE? Do you THRIVE in a FAST-PACED work environment? Do you wake every morning EXCITED to work with GREAT PEOPLE and create SUCCESS TOGETHER? Then Intermedia is the place for you.

Intermedia has established itself as a leading provider of cloud communications and collaboration tech that allows companies to connect better. We have a strong track record of growth, profitability, and creating an environment where everyone matters. Everyone. While we are fast-paced and admittedly a bit intense, we promise that you won’t be bored. You will find Intermedia is a place where you can indulge your passion for creating and supporting great cloud technology. What’s more, we always look to promote from within and have many employees who have been with us 10, 15, and 20+ years!

Culture at Intermedia is built on teamwork and transparency. We hold each other accountable and always have each other’s back!

While primarily remote, this role requires occasional visits to the office in Coimbra or in Aveiro. We plan to open an office in Porto in the future. This approach gives team members the flexibility to work remotely while also coming together in the office for collaboration and teamwork.

Are you ready to make your mark?

About the Role

We are looking for a Site Reliability Engineer (SRE) to improve the reliability of our AI and analytics platforms and the services that depend on them. As Intermedia expands its global product deployments, this role will strengthen SRE practices across production infrastructure, data pipelines, machine-learning services, customer-facing analytics, and AI-powered Voice and Unified Communications capabilities. You will partner with AI, data, product, and platform teams to build scalable, observable systems that deliver dependable insights and resilient customer experiences.

What you will be doing:

  • Run and improve production environments that support AI and AI workloads, data pipelines, analytics applications, and customer-facing services.
  • Build software and automation to manage cloud infrastructure, data platforms, model-serving infrastructure, and application services.
  • Define and measure service level indicators, service level objectives, and error budgets for AI and analytics services, including availability, data freshness, pipeline completion, and inference latency.
  • Build end-to-end observability that correlates metrics, logs, and traces with data-quality signals, AI-service performance, and customer impact.
  • Monitor and optimize the reliability, performance, capacity, and cost of batch and streaming workloads, analytics queries, and inference services.
  • Partner with data engineering and machine-learning teams to make ingestion, transformation, feature, training, deployment, and reporting workflows production-ready.
  • Automate CI/CD and production-readiness checks for data pipelines, model and prompt releases, schema changes, analytics applications, and dashboards.
  • Detect and resolve data-quality incidents involving missing, stale, delayed, or anomalous data, schema drift, and broken lineage or dependencies.
  • Design and test graceful degradation, dependency isolation, retry and fallback patterns, and recovery procedures for impaired AI, data, or downstream services.
  • Improve the reliability of AI-powered Voice and Unified Communications capabilities such as speech recognition, transcription, summarization, intelligent routing, conversational assistance, and text-to-speech.
  • Establish operational monitoring for model and AI-service behavior, including latency, throughput, error rates, drift indicators, and changes in output quality.
  • Plan capacity and run performance, load, and resilience tests across compute-intensive AI workloads, distributed data processing, and analytics services.
  • Lead incident response and post-incident improvement for AI and analytics services, using measurable actions to reduce recurrence and recovery time.
  • Support secure and reliable access to cloud storage, processing, and query services used by analytics products and internal decision-making.
  • Reduce operational toil and improve engineering productivity through platform tooling, runbooks, self-service automation, and clear operational standards.

What you will bring to the role:

  • Bachelor's degree in computer science, data engineering, software engineering, or another technical or scientific discipline, or equivalent practical experience.
  • 4-7 years of experience in production operations, systems engineering, SRE or DevOps, software deployment, and maintenance of distributed production systems.
  • Experience with cloud infrastructure, containers, Kubernetes, distributed systems, and scalable compute and storage services.
  • Experience operating data processing, orchestration, storage, or analytics technologies, such as Kafka, Spark, Airflow, dbt, data warehouses, or comparable cloud services.
  • Ability to use metrics, logs, traces, data-quality checks, freshness indicators, lineage, and service-level indicators to diagnose complex production issues.
  • Experience with CI/CD, DataOps or MLOps practices, automated testing, controlled rollout, and rollback of data and AI service changes.
  • Strong troubleshooting skills across Linux, applications, networks, APIs, data pipelines, and distributed service dependencies, with attention to security and access controls.
  • Strong analytical problem-solving and cross-functional communication skills, with a proactive approach to reliability, performance, and continuous improvement.

Diversity Inclusion and Equal Opportunity

We hire, promote, and compensate employees based on their ability to perform their job responsibilities, without regard to race, color, creed, religion, sex, gender, marital status, national origin, ancestry, age, citizenship, physical or mental disability, sexual orientation, or any other basis protected by applicable law (collectively referred to in our Code of Conduct as “Protected Classes”). We do not tolerate employment discrimination in the workplace, and we are committed to making reasonable accommodations for identified disabilities or other limitations as required by all applicable laws. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar jobs

Apply for this job