Cvent logo

Senior Site Reliability Engineer

Salary
$120K–$150K
USD per year
Hiring from
Canada, Spain, United States
Work type
Hybrid
Posted
Oct 2, 2026
Is this job info correct?

Overview

Our Culture and Impact

Cvent is a leading meetings, events, and hospitality technology provider with more than 6,000+ employees and 34,000+ customers worldwide, including 60% of the Fortune 500. Founded in 1999, Cvent delivers a comprehensive event marketing and management platform for marketers and event professionals and offers software solutions to hotels, special event venues and destinations to help them grow their group/MICE and corporate travel business. Our technology brings millions of people together at events around the world. In short, we’re transforming the meetings and events industry through innovative technology that powers the human connection.

Cvent's strength lies in its people, fostering a culture where everyone is encouraged to think like entrepreneurs, taking risks and making decisions confidently. We value diverse perspectives and celebrate differences, working together with colleagues and clients to build strong connections.

AI at Cvent: Leading the Future

Are you ready to shape the future of work at the intersection of human expertise and AI innovation? At Cvent, we’re committed to continuous learning and adaptation—AI isn’t just a tool for us, it’s part of our DNA. We’re looking for candidates who are eager to evolve alongside technology. If you love to experiment boldly, share your discoveries, and help define best practices for AI-augmented work, you’ll thrive here. Our team values professionals who thoughtfully integrate AI into their daily work, delivering exceptional results while relying on the human judgment and creativity that drive real innovation.

Throughout our interview process, you’ll have the chance to demonstrate how you use AI to learn, iterate, and amplify your impact. If you’re excited to be part of a team that’s leading the way in AI-powered collaboration, we’d love to meet you.

When a Fortune 500 company runs their annual sales conference, a software company broadcasts their product launch, or 500 exhibitors scan leads on the trade show floor — this team is why it works. We keep the infrastructure running for Cvent's event product suite, and we're rebuilding how SRE operates along the way.

In This Role, You Will:

Own reliability outcomes — not just tasks assigned to you.

At the Senior SRE level, you identify the work that matters, own it to completion, and understand the business context behind it. You'll be the embedded reliability partner for engineering teams across the portfolio — knowing each service's architecture, failure modes, and customer impact zones well enough to be genuinely useful, not just present. Deep product knowledge isn't a bonus; it's what separates great SRE work from good SRE work.

Co-own an active cloud-to-cloud migration without dropping production.

A meaningful portion of this team's current work is migration: moving search indexes, caches, and databases from Azure to AWS while keeping live traffic healthy. Senior SREs aren't bystanders to the cutover — you help design the validation gates, own rollback readiness, and know when to pause a migration before it becomes an incident. You're not inheriting a clean state. You're helping build one.

Drive SLOs and observability from service definitions, not just thresholds.

Some services have SLOs. Others don't have meaningful ones yet. You'll define SLIs that reflect what customers actually experience — search latency, availability, booking funnel health — instrument the journey, and build the alert posture that surfaces real problems instead of noise. We use Datadog APM and synthetic monitoring as the baseline. You'll make them better.

Make deployment reliable, not just frequent.

Our deployment surface includes CDK-managed stacks and Octopus Deploy pipelines across multiple environments. Deployment failures are a real source of incidents. You'll reduce that by improving health gates, validating releases before they reach production, and building rollback paths that work when you need them — not just in theory.

Mentor junior engineers and make the knowledge compound.

Senior SREs on this team review runbooks, lead incident debriefs, and help SRE I and SRE II engineers develop pattern recognition. The operational knowledge base — playbooks, postmortems, detection logic — is a team product, and you help build it. We share an on-call rotation across two sub-teams; your depth in one area makes the whole rotation stronger.

Here's What You Need:

AWS operational depth— ECS, Aurora/RDS, ElastiCache/Redis, OpenSearch at customer-serving scale; you've operated these, not just provisioned them

Database operations in production— Postgres or Aurora: connection management, query performance, migration safety, zero-downtime cutover patterns

Deployment orchestration— Octopus Deploy, Jenkins, or equivalent; you've debugged failing deploys, improved pipeline reliability, and owned rollback decisions under pressure

Scripting fluency— Python, Bash, or TypeScript; enough to automate your own toil and build the tooling your team actually uses

Observability tooling— Datadog at production depth: APM, monitors, SLOs, synthetic checks, alert tuning; you build the signal, you don't just read it

Incident command— you've led the call, written the postmortem, and driven root cause through to corrective action; stakeholders heard from you before they had to ask

SRE fundamentals— SLIs, SLOs, error budgets, toil accounting; not as buzzwords but as tools you've actually used to improve a system

Strong Plus

  • Search infrastructure — OpenSearch, Elasticsearch, or Azure Cognitive Search; index health, query performance, and migration patterns across search platforms
  • Azure-to-AWS migration experience — cloud-to-cloud with live traffic: DNS cutover, data sync validation, parallel-run validation gates, rollback sequencing
  • Daily use of Claude Code, Cursor, or AI coding assistants as an engineering accelerant — not occasional dabbling
  • CDK for infrastructure authoring — the migration workloads in this team are being built in CDK


Tech Stack AWS ECS - Aurora / RDS - Postgres - Redis / ElastiCache - OpenSearch - S3 - CDK - Octopus Deploy - CDF - Jenkins - CCID - Normandy Deployment - Patching Helper - Datadog - PagerDuty - Python - Bash - TypeScript - Claude Code

What we don't list but care about

  • You measure what matters and let data change your mind
  • You'd rather automate the fifth repetition than do it a sixth time manually
  • You're comfortable saying "I don't know yet" and uncomfortable with "that's just how it works"
  • You understand the difference between toil and real work, and you track which is which

Hybrid: 2 days in office

This job posting is intended to comply with all applicable laws. If we learn during the course of our recruitment process that, due to an applicant’s location, further information about the position is required, including certain salary information, this information in this posting will be supplemented accordingly.

The estimated base salary range for new hires into this role is $120,000 - $150,000 annually + bonus depending on factors such as job-related knowledge, relevant experience, and location. We also offer a competitive benefits package, details of which can be found here.

We are not able to offer sponsorship for this position

Physical Demands

LinkedIn Remote Type

Indeed Remote Type

WFH Flexible

Similar jobs

Apply for this job