DevOps Engineering Manager
- Hiring from
- United States
- Work type
- Remote
- Posted
508,412 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
About Logicbroker, Inc.
Logicbroker is the Agentic Commerce Orchestration Engine helping enterprise retailers, brands, suppliers, and distributors connect and grow. Our Intelligent Commerce Network powers over $10B in GMV for global leaders like Samsung, Walgreens, and Home Depot by automating the entire process from discovery to doorstep and stock to dock. We make products discoverable, shoppable, fulfillable, and returnable so our clients can grow faster, delight customers, and run smarter operations.
Job Summary
A DevOps Manager at Logicbroker owns the delivery, reliability, and health of the DevOps and Infrastructure team. The team runs the cloud platform under every Logicbroker and Virtualstock service. That covers Kubernetes, CI/CD, observability, and production environments serving more than 100 million requests each month.
Managers at Logicbroker are player-coaches. You answer for what your team ships, how reliably it runs, and how your engineers grow. You also contribute directly. Weight your own work toward investigation, runbooks, automation, and unblocking engineers. Leave critical-path project work to the team.
What You'll Do:
People and delivery
- Own delivery for the team. Turn roadmap priorities into a sequenced plan. Answer for what ships and when.
- Run hiring, onboarding, 1:1s, performance management, and career growth for every engineer on the team, including contractors.
- Shield the team's focus time. Absorb or redirect interrupt-driven requests instead of passing them to engineers.
- Report capacity honestly in planning. Push back on unrealistic asks instead of absorbing them silently.
- Act as first escalation point for delivery risk, interpersonal issues, and technical disagreements the team cannot resolve.
Reliability and operations
- Own production reliability. Define SLOs and SLIs for tier-1 services, report error-budget burn each month, and plan capacity ahead of demand.
- Contribute to the on-call rotation under the company incident and problem management policy. Set rotation size, handoffs, and escalation paths so that no engineer carries sustained out-of-hours load.
- Lead post-incident reviews for infrastructure causes. Track every corrective action to closure.
- Keep Datadog coverage and alert quality at a level where on-call engineers trust the pages they receive. Track and cut alert noise.
- Report cloud spend by service across AWS, Azure, and IBM Cloud, and act on the largest drivers.
Platform and migration
- Run the platform as a product. Treat application teams as your customers, publish a standard path for builds, deployments, secrets, and environments, and measure adoption.
- Keep production infrastructure changes in version-controlled code (infrastructure as code and GitOps).
Compliance and policy
- Own the infrastructure, access, change, and CI/CD policies assigned to your team in the company policy ownership matrix. Keep each one current, review it on its cycle, and map it to the relevant SOC 2 and ISO 27001:2022 controls.
- Check that the team's practice matches the written policy. Fix the practice, or change the policy through formal review. Record every exception with an owner and an expiry date.
- Produce audit evidence as part of normal work: access reviews, change and deployment records, patch status, backup and recovery tests, and incident records. Support the SOC 2 Type 2 and ISO 27001 surveillance audits, including auditor walkthroughs.
- Own the Cyber Essentials Plus technical controls for company infrastructure: boundary firewalls, secure configuration, malware protection, access control, and patching of hosts, operating systems, and container base images within the required window.
- Own the infrastructure security baseline for secrets and access control.
- Work with the Technical Operations Director, who coordinates audits, infosec documentation, and assessor relationships.
Engineering practice and leadership
- Set the team's standards for code review, testing, and change risk.
- Set the team's norms for AI-assisted development: what engineers verify, what needs elevated review, and where AI tools help or add risk.
- Own the team's contribution to the engineering metrics program (DORA or equivalent). Track the numbers, act on them, and answer for the trend.
- Join the engineering leadership on-call escalation path, not the primary rotation. Incidents reach you when they need a management decision on customer communication, resourcing, or cross-team coordination.
- Contribute to hiring plans, org design, and technical debt prioritization with the VP and peer Managers.
What We Need:
- 3+ years managing engineers, or equivalent experience as a tech lead with informal management scope moving into a formal Manager role.
- 5+ years in DevOps, site reliability, or infrastructure engineering, with current hands-on ability.
- Production Kubernetes experience. EKS is preferred.
- Ownership of CI/CD pipelines and infrastructure as code.
- Ownership of production reliability (on-call, incident response, SLOs and SLIs) for a system that handles 100M+ requests each month with off-peak batch workloads.
- Experience designing an on-call rotation that a small team can sustain.
- AI-assisted development as a standard part of your own work, and the judgment to set verification norms for your team
- A record of shipping on committed timelines without burning out the team
- Direct experience giving difficult feedback, managing underperformance, and making hiring and termination decisions
- Clear written and verbal communication. You translate technical risk into terms a VP or non-technical stakeholder can act on.
- Working knowledge of secure development practices and privacy law (CCPA, GDPR)
- Working knowledge of SOC 2, ISO 27001, and Cyber Essentials Plus requirements, and experience producing audit evidence for at least one of them
- Experience writing or maintaining engineering policies and procedures that people follow in practice
- Comfort making a call with incomplete information.
Preferred Skills:
- Experience managing a team through a cloud migration, merger, or major technology transition.
- Experience with Datadog or a comparable observability platform, including SLO-based alerting.
- Experience with ArgoCD, Istio, Crossplane, Karpenter, or comparable GitOps and platform tooling.
- Experience setting or evolving engineering metrics (DORA, SPACE, or similar).
- Experience managing full-time and contractor engineers across US and UK time zones.
- Experience leading a team through a SOC 2 Type 2 audit or an ISO 27001 surveillance audit.
- Experience with cloud cost management (FinOps).
- Enough familiarity with Python (Django), .NET Framework, or Go services to challenge deployment and runtime proposals from application teams.
Why Logicbroker:
Mission-Driven Culture: Be part of a company transforming digital commerce through innovation and agility—your work directly shapes how global brands connect with customers.
Collaborative, No-Ego Environment: We believe the best ideas win, not the loudest voices. You’ll work alongside teammates who challenge and support each other.
Flexibility with High-Performance Energy: Whether remote or in-office, we foster autonomy and accountability—because we trust you to own your success.
Leadership That Listens: Our executives are not just accessible—they’re invested in your growth, open to your ideas, and committed to building a company where people thrive.
Celebrated Wins, Shared Learnings: From team offsites to Slack shoutouts, we celebrate progress and learn from setbacks together