Platform Engineer
- Hiring from
- Canada
- Work type
- Remote
- Posted
- Oct 2, 2026
Senior Platform Engineer (SRE), AI Solutions
๐ Remote | Montreal - MUST be a Permanent Resume
Are you a senior-level Platform Engineer, SRE, or DevOps professional who thrives in complex, cloud-native production environments? Do you enjoy building the platforms that enable developers to move faster, operate reliably, and leverage AI-driven automation at scale? We have an immediate opening for a Senior Platform Engineer to help shape the next generation of AI-enabled software delivery and operations.
What You'll Do
- Design, build, and operate enterprise-grade Kubernetes platforms across AWS, Azure, and GCP.
- Create and automate cloud infrastructure using Terraform, Helm, GitOps, and Infrastructure as Code best practices.
- Own observability and monitoring platforms leveraging Prometheus and Grafana.
- Build reliable, scalable solutions supporting mission-critical applications, microservices, and high-volume production workloads.
- Partner with engineering, product, and business stakeholders to simplify complex technical concepts and communicate effectively with both technical and non-technical audiences.
- Lead reliability initiatives including incident response, performance optimization, SLOs/SLIs, automation, and operational excellence.
- Leverage AI and agentic technologies to automate monitoring, remediation, deployment, and operational workflows.
What We're Looking For
โ 10+ years of experience progressing through Cloud Engineering, DevOps, Platform Engineering, and/or Site Reliability Engineering roles
โ Deep hands-on Kubernetes experience, including designing, automating, maintaining, and scaling enterprise production clusters
โ Strong experience with AWS, Azure, and/or GCP in enterprise environments
โ Expertise with Terraform, Infrastructure as Code, CI/CD pipelines, and automation frameworks
โ Experience supporting highly available, large-scale systems consisting of multiple interconnected microservices
โ Strong background in monitoring, observability, performance tuning, and operational troubleshooting using tools such as Prometheus and Grafana
โ Experience working in complex, mission-critical environments where downtime, latency, security, and reliability are business-critical concerns
โ Excellent communication skills with the ability to explain technical concepts to business stakeholders, sales teams, legal teams, and non-technical users
#SRE #PlatformEngineer #DevOps #Kubernetes #Terraform #CloudEngineering #AWS #Azure #GCP #AI #AgenticAI #InfrastructureAsCode #SiteReliabilityEngineering #RemoteJobs #HiringNow #Syntax