Principal Dev-Ops Architect
CditsolutionsJob Description
This is a remote position.
Senior technical authority for a cloud platform running global, multi-tenant SaaS services. This is a hands-on individual-contributor architect role — not people management. The architect designs the platform, sets standards and reference implementations other teams build on, and still writes Terraform, builds CI/CD pipelines, and stands up the AI/ML platform personally. Defines how reliability is measured against SLOs, how releases ship, and how the AI/ML platform is built and governed while keeping the environment HIPAA-compliant and SOC 2 Type 2 audit-ready. Influences products, software, and QA through architecture and example.
Core Responsibilities
• Own platform architecture and technical roadmap for infrastructure, deployment, observability, and the AI/ML platform.
• Set engineering standards, patterns, and golden paths for Infrastructure as Code (IaC), CI/CD, and AI tooling; drive adoption through reference implementations and architecture reviews.
• Manage all cloud infrastructure as code in Terraform — reusable modules, remote state, peer-reviewed PRs, drift detection, and automated plan/apply in CI/CD.
• Enforce policy-as-code (OPA, Sentinel, or equivalent) so infrastructure changes meet security and cost guardrails before merge.
• Design and deploy AWS infrastructure across dev, UAT, staging, and production for performance, availability, recoverability, and security (CIS Critical Security Controls).
• Build and operate CI/CD pipelines for large-scale applications on AWS; own release management, rollback, blue/green, canary, and release gates.
• Package and run containerized workloads on Docker and Kubernetes (EKS).
• Lead the SLI/SLO/SLA program and modern observability using OpenTelemetry; drive down MTTD and MTTR; lead blameless post-incident reviews and participate in on-call.
• Provision and operate the AI/ML platform — Anthropic Claude via AWS Bedrock and internal MCP services — all managed as IaC.
• Build LLMOps practices: prompt versioning, evaluation pipelines, token cost attribution, guardrails, and audit logging of agent actions; enforce the PHI data boundary to BAA-covered providers only.
• Operate and evidence the platform controls required for SOC 2 Type 2 and HIPAA; own secrets management, supply-chain security (SBOM, image and dependency scanning), and FinOps.
Required Qualifications
• Bachelor's degree in Software Engineering or equivalent combination of technical education and work experience.
• 10+ years in SRE / DevOps / Platform Engineering delivering CI/CD, REST API deployment, containerization, IaaS/PaaS, data pipelines, and application observability — including time at a senior IC or architect level (Staff, Principal, or Architect).
• Proven technical authority across teams: sets architecture and standards and influences delivery through expertise and example rather than direct management.
• Demonstrated experience driving adoption of a new practice or platform (IaC, CI/CD overhaul, or an AI/ML platform) across multiple teams.
• Hands-on Terraform, including reusable modules other teams consume via self-service, remote state, and change management in a CI/CD pipeline.
• Building and operating CI/CD pipelines for large-scale applications on AWS (GitHub Actions, Jenkins, GitLab, or AWS-native).
• Running containerized workloads on Docker and Kubernetes.
• Monitoring and troubleshooting using cloud-native tooling and OpenTelemetry.
• Linux system administration, Unix scripting, and automation.
• Experience working in a HIPAA / HITECH / HITRUST / PHI / PII or PCI DSS environment.