ES

DevOps Engineer (ML Infrastructure)

Hiring from
Canada
Work type
Remote
Posted
Sep 29, 2026
Is this job info correct?

Role name: DevOps Engineer (ML Infrastructure)

Work Location: Canada (Remote)


DevOps Engineer (ML Infrastructure)

We are seeking a highly skilled Senior DevOps Engineer to help build, operate, and evolve large-scale, business-critical infrastructure platforms. This role focuses on reliability, scalability, automation, and operational excellence across distributed systems supporting high-volume production workloads.

You will work closely with software engineers, platform teams, and infrastructure specialists to deliver highly available services, improve developer productivity, modernize infrastructure, and drive innovation through automation and AI-assisted operations. This is an opportunity to solve complex technical challenges at scale while influencing the future direction of platform engineering and infrastructure management.

To design, build, and operate scalable machine learning infrastructure that enables efficient training, deployment, and monitoring of AI/ML workloads. The ideal candidate combines deep expertise in cloud-native technologies, Kubernetes, distributed systems, and software engineering to support large-scale machine learning platforms and GPU-based environments.

Key Responsibilities

  • Design, build, and maintain scalable MLOps platforms for training, deploying, and monitoring machine learning models.
  • Develop and manage cloud-native infrastructure supporting large-scale ML workloads on Kubernetes.
  • Implement and operate batch scheduling solutions such as Volcano and Kueue to optimize utilization of GPU and compute resources.
  • Build and maintain CI/CD pipelines for ML services, infrastructure, and platform components.
  • Automate infrastructure provisioning and lifecycle management using Infrastructure-as-Code practices.
  • Manage and optimize Kubernetes environments, including production workloads running on AWS EKS and other cloud platforms.
  • Lead and support infrastructure modernization and migration initiatives across cloud and platform ecosystems.
  • Partner with Data Scientists, ML Engineers, and Software Engineers to productionize machine learning solutions.
  • Implement robust observability, monitoring, and alerting for distributed systems and GPU clusters.
  • Ensure platform reliability, security, scalability, and operational excellence.
  • Troubleshoot complex distributed systems and performance bottlenecks across infrastructure and ML workloads.

Required Qualifications

  • 5+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or Infrastructure Engineering.
  • 3+ years of experience supporting production machine learning or AI platforms.
  • Strong experience with Kubernetes and large-scale workload orchestration.
  • Hands-on experience with Kubernetes batch schedulers such as Volcano / Kueue
  • Solid understanding of:
  • Distributed systems
  • Containerization technologies
  • Cloud-native architectures
  • Microservices-based platforms
  • Experience managing workloads on Kubernetes platforms such as AWS EKS.
  • Proven track record delivering and supporting infrastructure migration projects.
  • Strong programming skills in Python or Golang
  • Experience with Infrastructure-as-Code and deployment tools including Terraform / Helm
  • Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or Azure DevOps.
  • Strong Linux systems administration and troubleshooting skills.
  • Experience building observability solutions using:
  • Prometheus
  • Grafana
  • Cloud-native monitoring tools
  • Understanding of ML lifecycle management, model deployment, and production operations.
  • Preferred Qualifications
  • Experience supporting GPU-intensive machine learning or AI training platforms.
  • Hands-on experience with MLflow / Kubeflow
  • Experience with distributed training frameworks and GPU resource management.
  • Familiarity with LLMOps, Generative AI, RAG architectures, and vector databases is a plus
  • Knowledge of GitOps tools such as ArgoCD or Flux.
  • Experience with multi-cluster Kubernetes environments.Parquet



Mandatory skills:


  • Distributed systems
  • Containerization technologies
  • Cloud-native architectures
  • Microservices-based platforms CI/CD/Gitubs/Jenkkins,AWS EKS ML workflow Kubflow


Similar jobs

Apply on LinkedIn