COPE Health Solutions is hiring a Senior DevOps / Platform Infrastructure Engineer to help run and improve the platform behind our applications. Our infrastructure lives primarily in Microsoft Azure. We deploy through GitHub Actions, run workloads on Kubernetes, manage multiple databases, and support systems that handle healthcare data. This is a hands-on role. You'll spend more time in Bicep templates, AKS clusters, GitHub Actions workflows, PowerShell scripts, and architecture diagrams than you will in meetings about them. You'll also be one of the people developers call when a deployment breaks at 4:30 on a Friday afternoon. Patience, curiosity, and a healthy sense of humor go a long way. Because we work with healthcare data, security, compliance, and operational discipline aren't side projects. They're part of the job. FLSA Status Exempt Salary Range $128,600-$181,600 Reports To Director of Data Information and Security Direct Reports Yes Location Remote Travel Up to 10% Work Type Regular Schedule Full Time Position Description: Build and maintain Azure infrastructure using Bicep Operate and improve our AKS environments, including upgrades, scaling, reliability, and cost optimization Own CI/CD pipelines built with GitHub Actions Automate repetitive operational work using PowerShell and Python Implement monitoring, logging, and alerting that helps us discover problems before users do Reduce cloud spend without creating new operational risks Support PostgreSQL, SQL Server, and MySQL environments, including backups, maintenance, performance tuning, and disaster recovery planning Design and document network architecture, including hub-and-spoke VNets, private endpoints, subnet segmentation, peering, routing, and hybrid connectivity Support identity, access management, secrets management, and compliance initiatives Improve developer experience through better tooling, documentation, automation, and sensible platform defaults Participate in incident response, root-cause analysis, and operational reviews Documentation is part of the job We expect documentation to be treated like code. You'll be responsible for creating and maintaining: Runbooks Architecture diagrams Onboarding guides Incident postmortems Operational procedures Platform documentation If someone asks how traffic gets from Point A to Point B, there should be a diagram that answers the question. Qualifications: What we're looking for 5+ years operating production infrastructure Deep hands-on experience with Microsoft Azure Strong Infrastructure-as-Code experience with Bicep Extensive Networking experience Production Kubernetes experience, preferably AKS Strong GitHub Actions and CI/CD experience Strong PowerShell and Python scripting skills Experience supporting Linux and Windows Server environments Experience supporting PostgreSQL, and SQL Server Experience designing and troubleshooting network architectures Experience with observability platforms such as Azure Monitor, or similar tools Experience implementing RBAC, IAM, secrets management, and security controls Experience working in regulated environments such as HIPAA,HiTrust, and SOC 2 Strong written communication and documentation skills On-call and production support We run systems clients and their members depend on, so someone needs to be reachable when something breaks outside business hours. We try to be humane about how we do that. Platform engineers share a primary/secondary on-call rotation. Daytime issues get picked up by whoever's around; the rotation exists for evenings, weekends, and holidays. It covers genuine production incidents: a service down, a data-path failure, a security event, not routine tickets or "can you look at this sometime" requests. When you're paged, you're paged for something that actually matters. We expect platform engineers to participate in incident response, troubleshooting, and root-cause analysis. Just as importantly, we expect recurring operational pain to be addressed through automation, monitoring, documentation, or engineering improvements. The goal isn't to become better at responding to the same alert every week; it's to make sure that alert stops happening. We've all been on the receiving end of bad on-call rotations. We'd rather invest in reliability than heroics. What success looks like Month 1 Get access Learn the environment Ship something small Join incident reviews Figure out where the sharp edges are Month 2 Take ownership of one major platform area Contribute meaningful automation or operational improvements Start participating in production support activities Month 3 Identify risks we haven't seen yet Propose improvements we haven't thought of Help raise the engineering standard of the platform Interview process We're less interested in trivia than in how you think. Candidates should expect practical technical discussions, including: Reviewing a Bicep template Debugging a GitHub Actions pipeline Diagnosing an AKS issue Explaining a hub-and-spoke Azure network design Discussing a production incident they've personally handled Walking through a security or HIPAA-related scenario A strong answer doesn't require perfection. We're looking for engineers who can reason through problems, communicate clearly, and learn from operational experience. How we work You'll join a team of 13 and work closely with engineering and security teams. Code review is required. Ego is optional. Documentation gets reviewed like code. When incidents happen, we focus on fixing root causes rather than assigning blame. Benefits: As a firm passionate about health care, we’re deeply committed to the health and wellness of our own team members. We offer comprehensive, affordable insurance plans for our team and their families , and a host of other unique benefits, such as a yearly stipend for wellness-related activities, and a paid parental leave program. You can learn more about our benefits offerings here: https://copehealthsolutions.com/careers/ About COPE Health Solutions COPE Health Solutions is a national tech-enabled services firm powering success for health plans and for providers in risk arrangements. Our comprehensive NCQA certified population health management platform and highly experienced team brings deep expertise, experience, proven tools, and processes to improve financial performance and quality outcomes for all types of payers and providers. CHS de-risks the roadmap to advanced value-based payment and improves quality and financial performance for providers, health plans and self-insured employers. For more information, visit https://copehealthsolutions.com/about-us/ To Apply: To apply for this position or for more information about COPE Health Solutions, visit us at https://copehealthsolutions.com/careers/open-positions/
Software Engineer, Infrastructure
Gray Swan AI
Infrastructure Engineer
Appliedatomicsinc
Senior Infrastructure Engineer
Brook & Whittle
Sales Engineer Data Center & AI Infrastructure
Bold New Solutions - BNS Power
Senior Firmware Engineer - Development, Verification and Infrastructure
Nvidia
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Nvidia