We’re Heidi. We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries. Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart. We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again. If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi. The role This role sits in the core Platform/SRE team that owns production. You'll lead a small SRE team today, growing it as Heidi scales, while staying hands-on in incident response, on-call, system reliability, and day-to-day operations. We're looking for someone who has already built or scaled a reliability team, not someone stepping into management for the first time. You'll set the standard for how the team operates, hire and structure it as headcount grows, and represent SRE in conversations with engineering leadership. The role stays ops-heavy: you're expected to be in the systems yourself, not just running a roadmap from a distance. What you’ll do Participate in on-call and incident response. Respond to production incidents, contribute to service restoration, and keep communication clear while things are on fire. Lead incidents end-to-end, including the ones outside your immediate team. Improve operational reliability. Spot recurring issues and reliability risks, then drive fixes through better alerting, automation, system changes, or process improvements. Own the production environment. Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services across the team's remit. Strengthen observability. Build dashboards, alerts, logs, and traces that surface issues earlier and cut diagnosis time, with a bias toward signals people can actually act on. Reduce operational toil. Automate the repetitive stuff, simplify runbooks, and improve tooling so on-call and daily operations get easier and safer over time. Support safe change. Improve deployments, rollback mechanisms, and operational readiness so shipping changes doesn't mean rolling the dice on an incident. Contribute to operational practices. Write and maintain runbooks, run blameless post-mortems, and raise the bar on incident response as the team learns. Collaborate closely with engineers. Partner with product and feature teams on production readiness, service ownership, and what "reliable" should mean for their systems. Lead and grow the SRE team. Manage the team you inherit day to day, then hire and onboard as Heidi scales. Own on-call structure, career development, and team norms as headcount grows. Shape the team's direction. Decide where the team invests next, balance reliability work against team capacity, and represent SRE in planning with engineering leadership. What you'll need 7+ years in SRE, DevOps, platform, or operations-heavy engineering roles, including experience formally leading or managing a team. A track record of hiring, coaching, and growing engineers, not just leading incidents solo. Deep experience supporting production systems, including on-call rotations, and the credibility to still jump into an incident yourself. You're comfortable debugging live systems under pressure. Strong experience operating cloud infrastructure at scale (AWS preferred). Solid hands-on experience with Kubernetes and containerized workloads in production. Infrastructure as code experience (Terraform or similar). Hands-on experience with monitoring and alerting tools such as Datadog or Prometheus, including designing alerting strategy, not just consuming dashboards. Scripting or automation experience in Python, Bash, or similar. Hands-on experience defining and owning SLOs, error budgets, and capacity planning, not just familiarity with the concepts. Nice to have Experience scaling a team through fast headcount growth, not just steady-state management. Experience in regulated or security-sensitive environments. Familiarity with databases, queues, and caches in production. How we show up Build for the next decade, not next quarter . Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism. Lead, don't wait . We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that. Follow the evidence . Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford. Own the outcome . Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander. Ship, measure, go again . A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless. Live in clinicians' reality . Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it. Why Heidi? You’ll join a team focused on real-world impact over imaginary valuations and glossy PR. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders, and designers who’ve felt the moral and practical toll of what non-care feels like. True A-players progress extremely fast here. The nature of the scale-up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on outcomes > inputs, not process theatre. We all take the bins out, metaphorically and literally. Building what we’re building isn’t always easy. But we didn’t choose easy, we chose to build something that actually matters. We hold ourselves to a higher standard because healthcare demands it. If you join Heidi, you recognise that the deeper question isn’t whether AI can solve the global healthcare crisis, but whose hands will shape it. The work is hard, but you will trust and admire the people you work beside, and rest easy knowing you’re doing the defining work of your career. We take care of you. We offer a $1,000 annual learning and development budget, a $150/month health and wellness allowance, a $500 home office budget, 26 weeks paid primary parental leave and 18 weeks paid secondary parental leave, fertility support up to $10,000, four weeks of work from anywhere per year, and serious equity.
Senior Site Reliability Engineer- Remote
ClickHouse
Site Reliability Engineer (Australia, Remote)
Axon
Senior Site Reliability Engineer
Omilia
Machine Learning Engineer, Reliability
Fal Ai
Senior Site Reliability Engineer
Block
Site Reliability Engineer
Audinate