SR

Staff Platform Engineer

Sage Recruiting Inc.
Posted 4 hours ago
CanadaRemoteCA$210K–CA$260KEngineering & Development
Is this job info correct?

Staff Cloud Engineer

Remote (U.S. & Canada)

$210K–$260K base + equity


Sage Recruiting is partnering with an AI-native, early-stage infrastructure startup that's tackling a problem every engineering leader is quietly panicking about: AI coding agents now write code faster than any human team can review it, and the old ways of enforcing standards (wikis, checklists, manual review) were never built for that volume. Our client has built a guardrails engine that turns a company's engineering standards into automated enforcement, on every commit, every pull request, and every deploy, for human and AI-written code alike.


They're small, senior, and heavily AI-leveraged, moving with the speed and focus of a team several times their size. They've already landed their first paying customers and have real momentum building. Top-tier venture investors and an angel bench of well-known operators and creators from the developer tools world back them. This is a company at an inflection point between "early traction" and "real scale," actively going live with large enterprise customers, and this hire is one of the people who will build the infrastructure that the next chapter runs on.


The Role

This is a staff-level, deeply technical individual contributor role for a true backend/cloud/ops generalist. You'll be handed meaty, ambiguous problems and trusted to own them end-to-end: design, build, ship, and fix, with a seat in the on-call rotation like everyone else on the team. You'll write and own production Go code for the core platform, and you'll design and run the AWS and Kubernetes infrastructure it lives on, across both a hosted service and customer-managed deployments. A big part of the job is taking early, minimal product surfaces and making them enterprise-ready: authentication, backup, availability, and the reliability bar that large customers expect. You'll stay hands-on while helping decide what gets built and how, and you'll work directly with the company's forward-deployed engineers and customer platform teams on hard deployment and scale problems, feeding what you learn back into the product.


Who You Are:

  • You're a senior Platform Engineer who's comfortable working in code and cloud infrastructure.
  • You're comfortable being handed an ambiguous, high-stakes problem and can be trusted to run with it end-to-end without a fleshed-out playbook.
  • You think like a startup engineer: you know where the smart trade-offs are, where it's fine to cut a corner, and where it isn't, rather than defaulting to the most complete or "correct" solution.
  • You lean Kubernetes-strong in particular.
  • You're AI-native in your daily workflow, and you're excited by the idea of taking a product that's complete but still early and making it ready for large enterprise customers.


What You'll Do

  • Write and maintain production Go code in the backend and supporting services, owning features from design and testing through deployment and ongoing operation
  • Design and evolve the AWS architecture and Kubernetes infrastructure, making deliberate trade-offs around reliability, security, performance, cost, and how much complexity a small team can support
  • Build reusable Terraform modules, Helm charts, and deployment tooling that make provisioning, upgrades, and recovery repeatable, for the hosted service and customer-managed installations alike
  • Harden early-stage product surfaces to meet enterprise expectations: authentication, backups, availability, and the operational maturity large customers require
  • Improve production reliability: useful metrics and alerts, capacity planning, backups and tested recovery, safe rollouts, and clear rollback paths. Take your turn in the on-call rotation and fix the root causes of recurring problems, not just the symptoms
  • Debug across the full stack, following a failure through Go code, a database query, container behaviour, Kubernetes networking, or cloud infrastructure rather than stopping at a team boundary
  • Improve CI/CD and the development environment so engineers can test realistic changes and ship frequently without making production fragile
  • Partner with forward deployed engineers and customer platform teams on difficult deployment and scale problems, feeding what's learned back into the product
  • Lead technical design and code reviews, mentor teammates, and document decisions well enough that other engineers can maintain what you build
  • Build and manage AI-assisted engineering workflows for implementation, review, testing, and operational investigation, with appropriate access controls and checks on their output


You must have:

  • Strong Go and software engineering fundamentals: production services, concurrency, APIs, testing, performance work, and debugging distributed systems. This doesn't need to be your single deepest specialty, but you're comfortable owning backend code end-to-end
  • Deep cloud experience, ideally AWS (we're open to strong GCP or Azure backgrounds too): networking, IAM, compute, storage, and managed databases, with a real understanding of failure modes, isolation boundaries, and cost
  • Extensive, strong Kubernetes experience: built and operated production clusters and workloads, handled upgrades, debugged real failures, and worked with scheduling, networking, storage, RBAC, resource management, and Helm beyond an install guide. This is where we most want depth
  • Extensive Terraform experience: reusable modules, state, environment separation, drift, and safe changes to existing production infrastructure
  • Strong operational instincts: comfortable with Linux, Docker, networking, and troubleshooting under pressure, and you've owned systems after launch and made them easier to operate over time
  • Working knowledge of production data stores and observability: you can operate a SQL database (PostgreSQL or comparable) and object storage like S3, and you know how to instrument and read metrics, logs, and traces (Prometheus, Grafana, OpenTelemetry) well enough to actually run a system, not just build one
  • Extensive, hands-on use of AI coding tools and agents as part of your daily work. You build your own workflows, supply the context and tools they need, manage their permissions/cost/failure modes, and can explain how you verify what they produce
  • A track record of owning ambiguous, critical problems end-to-end and delivering changes that held up in production, whether that's 7+ years as a staff-level IC or a faster growth trajectory that's gotten you there
  • Clear communication and high ownership: you write useful design notes, can explain a trade-off to another engineer or a customer, and make progress without waiting for a detailed spec


  • NOTE: This isn't a fit for someone who wants an architecture or management role with little hands-on coding, or who enjoys operating infrastructure but doesn't want to build production software. It's also not a fit for someone who uses AI only occasionally or needs a narrowly defined remit, with a separate team handling everything outside it.


It's a bonus if you have:

  • Something that makes you stand out: a well-known open-source project you've built or maintained with real traction, ownership of a critical production system early in your career, multi-disciplinary range (ops plus frontend, for example), or a background at a company known for excellent engineering craft
  • Experience building developer tools, infrastructure products, or other B2B software used by engineering teams
  • Experience shipping software into customer-managed Kubernetes environments, including restricted networks and enterprise security requirements
  • gRPC and Protocol Buffers, queues/background workers, or distributed job execution
  • GitOps and progressive delivery, Kubernetes operators, or multi-region systems
  • Familiarity with software supply-chain security, policy as code, or compliance automation
  • A founder or early-stage startup background, or a mix of big-tech and startup experience


What Success Looks Like

  • First 30 days: you're getting familiar with as much of the system as possible, shipping small, complete contributions across most areas of the product to build context fast.
  • By 60 days: you're taking on larger, more substantial projects.
  • By 90 days: you're consistently taking on real work and shipping it end-to-end without oversight.
  • Six months in: you're a trusted, independent owner of critical parts of the system, someone the rest of the team can hand hard problems to.


Why Join

  • Get in early with a company at a genuine inflection point, actively going live with large enterprise customers and scaling fast
  • Every team member has a real voice in product direction and execution, not just their own corner of the codebase
  • You'll be supporting the developer infrastructure of large enterprises: your work directly shapes the experience of other engineers, including platform teams at big, complex organizations
  • Real architectural ownership from day one, working directly with the founder and founding engineers, with room to shape both the product and how the team builds it
  • A small, senior team moving fast, with heavy investment in AI tooling and none of the bureaucracy of a bigger company
  • Growth here is squarely on the IC track: real room to grow your technical scope and seniority
  • Fully remote, with the team kept within a few hours of each other, so there are no late-night meetings


Compensation & Benefits

$210,000–$260,000 base plus equity

401(k)/RRSP matching

Healthcare, dental, and vision, including dependents

Life insurance

Health/Lifestyle stipend

Fully remote work


Similar jobs