Oliver James logo

Senior Site Reliability Engineer - AI-focused

Oliver James
Posted 3 hours ago
IrelandHybridEngineering & Development
Is this job info correct?
  • 12-month contract (potential to go permanent)
  • Hybrid working model (with occasional shift work)
  • Office location in Dublin
  • Day-rate contract

Overview

We are looking for an AI-focused SRE professional centred on the running, deployment and operational support of AI workloads and internal platforms. You will work closely with developers to build, deploy and operate internal AI-enabled tools and services, while ensuring they remain reliable, observable and scalable across production and near-production environments.

As an SRE, you will be a steward of production, with a strong focus on uptime, observability, operational resilience and continuous improvement.

Top 3 Must-Have Requirements

  1. Splunk & Dynatrace
    Hands-on experience with Splunk and Dynatrace, including monitoring, developing queries, creating alerts and supporting production environments.
  2. AI Workloads
    Experience running, deploying or supporting AI/ML workloads, preferably in production or near-production environments.
  3. CI/CD & DevOps Automation
    Strong experience with CI/CD and deployment tooling, including Git/Bitbucket, Jenkins, Maven, Artifactory, Chef and/or XLR.

Key Responsibilities

  • Act as a steward of production, ensuring services remain reliable, available, observable and operationally resilient.
  • Support the full lifecycle of services from design and development through deployment, operation and continuous improvement.
  • Monitor and maintain production and near-production environments, focusing on uptime, observability and system health.
  • Develop and improve Splunk and Dynatrace queries, dashboards, alerts and monitoring capabilities.
  • Investigate production incidents and use observability data to identify root causes and improve reliability.
  • Support the running and deployment of AI workloads and internal AI-enabled services.
  • Work closely with developers to develop, deploy and operate internal AI tools and platforms.
  • Support and improve CI/CD pipelines and deployment processes.
  • Run deployment plans using tools such as XLR and identify opportunities to improve reliability, efficiency and automation.
  • Develop automation to reduce manual intervention and improve software delivery.
  • Review operational and deployment processes and recommend improvements.
  • Support services before launch through system design reviews, capacity planning and operational readiness assessments.
  • Measure and monitor availability, latency, performance and overall system health.
  • Scale systems sustainably through automation and engineering best practices.
  • Analyse IT service management activities and provide feedback on operational gaps and resiliency concerns.
  • Participate in incident response and blameless post-incident reviews.
  • Troubleshoot issues across the technology stack to reduce mean time to recovery (MTTR).
  • Work collaboratively with development, operations and product teams.
  • Share technical knowledge and mentor junior engineers.
  • Collaborate with globally distributed teams across multiple locations and time zones.

Experience & Qualifications

Essential

  • 3-5 years' experience in SRE, DevOps, Production Engineering, Platform Engineering or a related discipline.
  • Strong hands-on experience with Splunk and Dynatrace, including monitoring, observability, queries and alerting in production environments
  • Bachelor's degree in Computer Science, Engineering, Mathematics, Physics or a related technical discipline, or equivalent practical experience.
  • Relevant cloud, DevOps, SRE or IT service management certifications are desirable.
  • Strong experience with CI/CD and DevOps automation, including Git/Bitbucket, Jenkins, Maven, Artifactory, Chef and/or XLR.
  • Experience supporting software deployments and production operations, with a focus on reliability, availability and incident response.
  • Experience running, deploying or supporting AI/ML workloads, ideally in production or near-production environments.
  • Strong scripting/programming skills, ideally Python and Shell.
  • Experience troubleshooting complex issues across the technology stack.
  • Strong problem-solving, communication and stakeholder management skills, with the ability to work effectively across development, operations and product teams.

Desirable

  • Fintech or financial services experience, particularly within highly available, business-critical environments.
  • Experience with AI platforms, data platforms or large-scale distributed systems.
  • Experience with Apache NiFi and SQL.
  • Additional programming experience with Java, C, C++, Go, Perl or Ruby.
  • Experience mentoring engineers or providing technical leadership.

What Success Looks Like

Success will be measured through improvements in uptime, observability, deployment reliability, incident response and MTTR, alongside increased automation and reduced manual operational effort.

Key Notes

Role focus: SRE, Production Operations, AI Workloads and DevOps Automation
Core technologies: Splunk, Dynatrace, Git/Bitbucket, Jenkins, Maven, Artifactory, Chef, XLR
Industry: Fintech/Financial Services preferred

Similar jobs