ContactMonkey logo

Intermediate Site Reliability Engineer

Salary
CA$130Kโ€“CA$150K
Hiring from
Canada
Work type
Hybrid
Posted
Is this job info correct?

513,132 remote jobs, straight from company career pages

100% free ยท New jobs every hour

Show job description

Hey there! We're ContactMonkey ๐Ÿ‘‹

Our mission is to power measurable employee engagement worldwide, and we're looking for an Intermediate Site Reliability Engineer to join our Engineering team.

About the job

We're looking for someone who enjoys running production systems, understands how applications work, and brings practical application security experience.

You'll work closely with our SRE and development teams to maintain our infrastructure, improve deployments, investigate production issues, and address security risks. You'll also contribute to the technical controls that support SOC 2 audits and GDPR compliance.

Our environment includes AWS, Kubernetes on EKS, Terraform, Terragrunt, GitHub Actions, Prometheus, Grafana, and CloudWatch. Our applications use Ruby on Rails, Vue.js, and Node.js, with MySQL, PostgreSQL, and Sidekiq supporting the backend.

You'll take ownership of defined projects and operational improvements, with senior engineers available for guidance and review. There's room to develop deeper expertise in reliability, infrastructure automation, and application security.

Your impact

  • Infrastructure & reliability: Maintain AWS and Kubernetes environments, troubleshoot production issues, and improve availability, performance, and resource usage.
  • Terraform & Terragrunt: Build and maintain infrastructure as code, review plans, manage environment configuration, and address infrastructure drift.
  • Deployments & developer experience: Improve CI/CD pipelines, deployment automation, release checks, and rollback procedures.
  • Monitoring & incidents: Improve monitoring and alerts, join the on-call rotation, and contribute to incident reviews.
  • Application security: Work with developers to assess vulnerabilities, review security risks, and validate fixes.
  • CI/CD security: Maintain code, dependency, secrets, container, and infrastructure scanning.
  • Cloud security: Strengthen IAM, secrets management, network controls, and Kubernetes security.
  • SOC 2 & GDPR: Support technical controls, audit evidence, and personal data protection.
  • Recovery: Test backups and recovery procedures, maintain runbooks, and support production readiness.
  • Collaboration: Participate in code reviews, document changes, and support application, AI, and data engineering teams.

About you

  • Around 3โ€“5 years of experience in SRE, DevOps, platform engineering, cloud operations, or a related engineering role. Equivalent practical experience is welcome.
  • Hands-on experience supporting production workloads in AWS.
  • Experience writing and maintaining Terraform modules and Terragrunt configuration, including reviewing plans and working with remote state.
  • Experience with Docker and Kubernetes, including troubleshooting deployments, services, health checks, and resource limits.
  • A solid understanding of Linux, networking, DNS, HTTP, and TLS.
  • Ability to write automation in Python, Bash, Ruby, JavaScript, or another suitable language.
  • Experience with Git, pull requests, CI/CD pipelines, and deployment workflows.
  • Experience using logs, metrics, dashboards, and alerts to investigate production problems.
  • Practical application security experience through vulnerability remediation, secure code review, threat modelling, or security tooling.
  • Understanding of common web application risks, including broken access control, injection, authentication weaknesses, and sensitive data exposure.
  • Familiarity with IAM, least privilege, secrets management, and encryption.
  • Working knowledge of SOC 2 controls and GDPR principles relevant to engineering, including access restrictions, data minimization, retention, and deletion.
  • Clear communication skills and good judgment about when to work independently, request a review, or escalate an issue.

How you can stand out

  • Supported Ruby on Rails or Node.js applications in production
  • Worked with MySQL, PostgreSQL, Redis or Valkey, and Sidekiq
  • Experience with GitHub Actions, Argo CD, Helm, Karpenter, or KEDA
  • Built useful monitoring with Prometheus, Grafana, or CloudWatch
  • Supported services across multiple AWS regions
  • Helped remediate penetration-test findings or contributed technical evidence to a SOC 2 audit
  • Participated in backup restoration or disaster recovery exercises.
  • Familiar with securing AI integrations, agent workloads, or MCP services
  • You hold relevant AWS, Kubernetes, Terraform, or security certifications.

Working with the team


You'll work closely with SRE and application engineers, and support AI and data initiatives where they depend on shared infrastructure.

We value people who ask questions, explain their reasoning, and leave systems easier for others to understand. You'll take part in technical discussions and code reviews, share what you learn, and help improve how we operate.

Reducing recurring incidents, unnecessary alerts, and repetitive manual work is part of the role. We want reliable systems and sustainable operations for the people supporting them.

What we bring to the table


๐Ÿฅ 100% employer-paid benefits and a Health Spending Account from day one
๐ŸŒŽ Work from anywhere in the world for up to four weeks
๐Ÿ’ฐ A stock option plan so you can own a piece of our success
๐Ÿ’ฒ An RRSP Group Savings Plan
๐Ÿ A generous vacation package
๐Ÿ“š A personal development budget
๐Ÿง– One personal day and two volunteering days
๐ŸŽ‚ Your birthday off
๐ŸŽ Five health days per year
๐Ÿ’ผ A downtown Toronto office for hybrid work, with plenty of snacks

Compensation and work details

The salary range for this role is $130,000-$150,000 Compensation is based on experience, skills, and our internal compensation framework and equity.

We're happy to discuss compensation throughout the hiring process.

This is a full-time position on our SRE team.

The role includes a shared on-call rotation after onboarding. We'll discuss the schedule, escalation support, and expectations during the interview process.

Who we are


ContactMonkey helps organizations create, send, and measure internal communications directly within Outlook and Gmail.

Our platform brings together email design, employee engagement tools, and analytics so internal communications teams can understand what reaches their people and what gets a response.

As the product grows, we're investing in reliability, security, and tooling that helps our engineering teams deliver changes confidently.

Diversity is our strength
At ContactMonkey, we're building products for diverse organizations, and we need a diverse team to do that. We strongly encourage applications from everyone regardless of race, religion, colour, national origin, gender, sexual orientation, age, marital status, or disability status.

We are committed to an accessible hiring process. If you need accommodations or adjustments during interviews or beyond, please let us know so we can arrange the support you need.

AI Disclosure
We use AI to take notes during our interviews. Applications and interviews are reviewed by our Talent Acquisition team. Our applicant tracking system uses AI for workflows and hiring process efficiencies.

Similar jobs

Apply for this job