Senior Software Engineer, DevOps & Security
ToastJob Description
Toast is recruiting on behalf of a fast growing B2B SaaS company in the competitive intelligence space. The company uses AI to turn market, competitor, and buyer signals into insights that revenue teams can act on, helping organizations win more competitive deals. Based in Vancouver with hubs in other major cities, the team is small, globally distributed, and scaling quickly, having grown through several recent acquisitions. They are hiring a Senior Software Engineer, DevOps & Security to make the right way to build and ship the easy way for every engineer, working alongside the DevOps team, engineering managers across Product, Data and AI, and Security and Compliance stakeholders. The team works office-first with a hybrid touch, together in person on Mondays, Wednesdays and Thursdays.
Responsibilities
Toast is recruiting on behalf of a fast growing B2B SaaS company in the competitive intelligence space. The company uses AI to turn market, competitor, and buyer signals into insights that revenue teams can act on, helping organizations win more competitive deals. Based in Vancouver with hubs in other major cities, the team is small, globally distributed, and scaling quickly, having grown through several recent acquisitions. They are hiring a Senior Software Engineer, DevOps & Security to make the right way to build and ship the easy way for every engineer, working alongside the DevOps team, engineering managers across Product, Data and AI, and Security and Compliance stakeholders. The team works office-first with a hybrid touch, together in person on Mondays, Wednesdays and Thursdays.
Responsibilities
- Own the paved road for infrastructure by building and extending a Pulumi component library so teams get correct, disaster-recovery-ready infrastructure by default across GCP environments
- Run GKE like a product, handling cluster and node pool upgrades, in-cluster services, workload right-sizing and policy enforcement as the company scales
- Own the core building blocks of CI/CD workflows, cut CI time and flake rate across the fleet, and standardize the golden path for shipping services
- Improve internal developer platform surfaces so engineers can see what is happening in their own systems
- Own the health of the observability stack, including consistent instrumentation, actionable alerts and less on-call noise
- Drive GCP and LLM cost reduction through attribution, right-sizing, commitment planning and removal of orphaned resources
- Extend software supply chain controls, own infrastructure identity and access, and build SOC 2 and other compliance requirements into automated, evidenced controls
- Balance proactive and reactive work, jumping in when the platform is on fire and always driving the change that stops the next incident
- Automate recurring manual work, and turn incident reviews into durable platform fixes
- Publish SLOs and error budgets for the services that matter most
- Deep, hands-on experience in DevOps, SRE, platform or infrastructure engineering, operating production systems that paying customers depend on, including being on-call for them
- Deep, practical Kubernetes knowledge
- Infrastructure as code treated as software, using Pulumi or Terraform
- Strong GCP or AWS fundamentals
- A background in systems engineering and writing software
- Experience owning CI/CD as a product, building and maintaining pipelines used by other teams (GitHub Actions or Buildkite preferred, GitLab CI or CircleCI experience transfers)
- Hands-on experience with security tooling as an engineering practice
- Compliance engineering experience for SOC 2 or enterprise partner security programs, implemented as automated controls, plus threat modeling
- Ownership under ambiguity: you scope a vague problem, ship it incrementally and clearly communicate what changed and why
- Fluency with AI coding tools and experience working on AI-first teams
- Experience owning platform migrations and decommissions would be a plus
- Having built an internal developer platform or self-service tooling that other engineers adopted is considered an asset
- Comfort with LLM infrastructure, such as gateways, token cost attribution, rate limiting and provider failover, would be a plus
- Tech stack includes GCP, Kubernetes, Pulumi, Postgres and Temporal
- Extended health and dental coverage that starts on day 1
- Eligibility for the Employee Stock Option Plan for all full-time employees
- Take-the-time-you-need vacation policy
- Mac or PC of your choice and access to top-tier tooling
- Fitness and lifestyle perk discounts
- Performance rewards, career development, coaching and annual performance reviews
- Annual company-wide kickoff in Vancouver, regular hub social events, and dog-friendly offices in Vancouver and Toronto
- Direct access to the leadership team, including the CEO
- Own real infrastructure at real scale, with authority over the IaC, GKE and CI/CD architecture every engineering team depends on
- Make an immediate, measurable difference: faster pipelines, fewer flaky failures, lower GCP and LLM spend with clear per-team attribution, and on-call page volume that drops every quarter
- Turn SOC 2 and audit readiness into automated controls, with audit findings in your areas trending to zero
- Help shape how a fast-moving, AI-first engineering organization ships software safely
- Application Review
- Vetting Call
- Profile Creation
- Client Submission