Engineering Manager, Production Platform and Orchestration
- Salary
- $200K–$300K
- Hiring from
- United States
- Work type
- Remote
- Posted
510,258 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
About Positron AI
Positron AI is building next-generation AI inference accelerators designed from the ground up for low-latency, high-throughput large language model inference. Our first-generation ASIC, Asimov, is a cutting-edge accelerator targeting frontier AI workloads, with additional generations already underway.
Role Overview
Positron is seeking an Engineering Manager to lead Production Platform and Orchestration within our Upstack Engineering organization. This team owns the software and operating practices that provision, deploy, observe, upgrade, and reliably operate Positron systems in production. You will inherit a technically strong core team and help it grow into a durable organization capable of supporting a fleet that is expanding by several multiples.
This is a technical leadership role with real operational accountability. You will set direction, build the team, create clear ownership, and improve the systems and processes behind fleet orchestration, deployment lifecycle, observability, incident response, release automation, and production reliability. You will work closely with serving and API, model enablement, compiler and runtime, hardware, customer-facing, and data center partners.
The strongest candidate will combine systems depth with organizational judgment, moving comfortably between architecture, delivery, incidents, people development, and cross-functional planning. This description intentionally emphasizes outcomes and ownership over a fixed organizational chart. As the fleet and customer base grow, the function may develop dedicated groups for fleet orchestration and capacity, deployment lifecycle, reliability and observability, data center operations, customer production operations, and operational tooling.
Key Responsibilities
Team Leadership and Organizational Growth
- Lead, coach, and grow a team of engineers spanning production platform, fleet orchestration, reliability, and operational automation.
- Hire thoughtfully, develop emerging leaders, and create ownership boundaries that remain effective as the organization scales.
- Translate customer and business priorities into sequenced engineering work while protecting the team from reactive, unstructured operations.
Platform Strategy and Roadmap
- Establish a clear technical and organizational roadmap for provisioning, deployment, environment lifecycle, fleet health, capacity, upgrades, and rollback.
- Build reliable orchestration and control-plane capabilities for inventory, placement, configuration, health management, and multi-system operations.
- Prepare the production platform for a heterogeneous accelerator environment that may include FPGA, ASIC, and GPU infrastructure.
Reliability and Production Operations
- Define service-level objectives, operational metrics, alerting standards, and a sustainable on-call model for customer-facing production systems.
- Own the operating cadence for incidents, escalations, postmortems, corrective actions, launch readiness, and reliability reviews.
- Increase automation across deployment, upgrades, remediation, diagnostics, capacity planning, and common support workflows so that fleet growth does not require linear headcount growth.
Cross-Functional Partnership
- Partner with hardware and data center teams on rack bring-up, networking, firmware, sparing, failure handling, and platform transitions.
- Collaborate with Distributed Serving and API, Model Enablement, and Compiler and Executor teams to turn new capabilities into supportable production services.
Required Qualifications
- Demonstrated success managing and growing engineering teams responsible for distributed systems, cloud infrastructure, production platforms, SRE, or a closely related domain.
- Strong technical judgment across Linux systems, networking, orchestration, deployment systems, observability, and production reliability.
- A proven record of turning ambiguous operational demands into a coherent roadmap, explicit ownership, and measurable engineering outcomes.
- Ability to recruit, coach, and retain engineers across experience levels while maintaining a high technical bar.
- Comfort operating across software, hardware, data center, customer, and business boundaries.
- Excellent written and verbal communication skills, with sound prioritization and the ability to make tradeoffs visible to technical and executive stakeholders.
- A hands-on leadership style, close enough to architecture and operations to ask the right questions without becoming the team's bottleneck.
Preferred Qualifications
- Hands-on experience operating GPU, FPGA, ASIC, or other accelerator fleets in production.
- Ownership of services with meaningful availability expectations, including on-call, incident management, root-cause analysis, and reliability planning.
- Experience building control planes, schedulers, placement systems, capacity-management systems, or multi-rack orchestration.
- Background in data center deployment, hardware lifecycle, firmware coordination, sparing and RMA processes, or production networking.
- A track record of scaling infrastructure from early deployments to multiple sites or hundreds of systems.
- Exposure to customer-facing infrastructure where engineering teams participate in production escalation and service readiness.
- Demonstrated automation work that materially reduced operational toil, incident frequency, or recovery time.
Leveling & Scope
While this role is currently posted at a specific level, we are a growth-oriented organization and are open to hiring at a more senior level for the right candidate. Please note that this job description serves as a focused but generalized overview of the role; specific responsibilities and impact expectations will be tailored to the experience and seniority of the final hire.
What Success Looks Like
In the first six months, you will build trust with the team and partner organizations, clarify ownership, decision rights, and the near-term hiring plan, and baseline fleet health, incident load, deployment reliability, operational toil, and the largest single points of failure. You will establish a practical operating cadence for on-call, incident review, release readiness, and reliability prioritization, and produce an agreed roadmap that balances immediate production needs with platform investments and automation. Between six and twelve months, you will grow the team and create durable ownership for fleet orchestration, deployment lifecycle, observability, and production reliability, while improving automated provisioning, upgrades, rollback, health monitoring, and operational diagnostics. Avoidable incidents and manual intervention will decrease, deployment confidence and customer readiness will increase, and you will be developing engineers and technical leads who can independently own major platform and operational domains. By twelve to eighteen months, you will be operating a resilient production organization with clear specialties, healthy management span, and sustainable coverage, supporting a substantially larger and more diverse fleet without proportional growth in operational effort, and demonstrating measurable improvement in availability, deployment speed, upgrade safety, incident recovery, and operational efficiency.
Why Join Us?
- You will build the production platform that turns purpose-built inference silicon into services customers can depend on, with direct ownership of how a rapidly growing fleet is deployed, operated, and scaled.
- You will shape both the technology and the organization from an early stage, defining the orchestration, reliability, and automation foundations that Positron will operate on for years to come.
Compensation & Benefits
The base salary range for this role is $200,000 – $300,000.
Please note that the figures provided represent the base salary range only and do not include other elements of our total compensation package, equity, or comprehensive benefits.
At Positron AI, we value the unique expertise each candidate brings. While the range above reflects our typical expectation for the position, we reserve the flexibility to exceed this range for candidates whose specialized skills, significant experience, or unique qualifications fall outside the standard scope of the role. Final offers are determined based on a variety of factors, including internal equity, and individual impact.
Benefits & Perks
We want you to do your best work and feel confident that you and your family are taken care of. That means comprehensive coverage, real time to rest, and support for your future.
Health and wellness
- Fully company-paid medical, dental, and vision insurance for you and your dependents
- Company-paid life and disability coverage, with voluntary options to add more
- Supplemental hospital, critical illness, and accident coverage available
Time off and flexibility
- Unlimited paid time off, we encourage everyone to truly unplug and recharge
- 13 paid company holidays
- Remote-first culture with a company-provided computer and home office setup
Compensation and future
- Competitive salary and equity
- 401(k) with company matching, eligible from day one
Visa Support
This position is open to candidates currently authorized to work in the U.S. We cannot provide new visa sponsorship for this role but are open to facilitating H-1B visa transfers for eligible candidates.
Equal Opportunity Employer. If you're excited about the role but don't meet every bullet, we'd still love to hear from you.