About Gray Swan
Gray Swan is on a mission to empower the world to use AI safely and securely. We evaluate AI models for the leading frontier labs along with building real-time threat detection and adaptive adversarial red teaming agents for teams deploying AI.
We're a team of approximately 60 people, well-funded, growing quickly. Our work directly influences how the world deploys AI agents and systems at scale.
The Role
Gray Swan is looking for an Engineering Manager, Enterprise to lead the team responsible for building and scaling the systems that power our AI security platform. You'll combine strong technical judgment with people leadership to help design the enterprise experience, including infrastructure, backend services, and cloud architecture that enable Gray Swan to build AI systems and for customers to safely deploy frontier AI models at scale.
This role is ideal for an engineering leader who enjoys solving complex infrastructure challenges while building and developing a high-performing team. You'll work closely with machine learning engineers, product engineers, and security researchers to ensure our platform remains fast, resilient, secure, and scalable as we grow.
You'll have significant ownership over foundational systems, technical strategy, and team development. As one of the leaders shaping Gray Swan's infrastructure organization, you'll have the opportunity to influence architecture, engineering practices, and how we build and operate systems in a rapidly evolving AI startup.
What You’ll Do:
Lead, mentor, and develop a team of infrastructure engineers, creating a culture of ownership, technical excellence, and continuous improvement.
Set technical direction and priorities for enterprise solutions in partnership with engineering leadership and other technical teams.
Oversee the design, development, and operation of highly available backend services and distributed systems that power Gray Swan's AI security platform.
Guide the team in owning and scaling cloud infrastructure across Kubernetes, AWS, GCP, Azure, networking, storage, and compute.
Drive the development of internal platform services, infrastructure tooling, and systems that improve developer productivity and reliability.
Establish and improve engineering practices around observability, monitoring, incident response, deployment, and operational excellence.
Partner closely with machine learning, security, and product engineering teams to build infrastructure capable of supporting increasingly complex AI workloads.
Identify and address opportunities to improve system performance, reliability, scalability, and infrastructure efficiency.
Help define long-term architecture and technical strategy while balancing immediate product and engineering priorities.
Participate in hiring and help build a world-class infrastructure engineering organization.
Provide technical guidance during complex production incidents and ensure the team learns from failures through effective post-incident reviews.
Communicate infrastructure strategy, tradeoffs, risks, and priorities clearly across engineering and company leadership.
Who You Are:
7+ years of experience building enterprise solutions with experience implementing backend infrastructure or distributed systems in production environments.
2+ years of experience managing, mentoring, or leading software engineering teams.
Strong technical background in infrastructure, distributed systems, backend engineering, or platform engineering.
Strong programming experience in C/C++, Go, Python, Rust, Java, or a similar language.
Experience operating production services on Kubernetes and modern cloud platforms such as AWS, GCP, and Azure.
Deep understanding of networking, distributed systems, containers, service orchestration, and scalable architectures.
Experience designing and operating APIs, microservices, asynchronous systems, and event-driven architectures.
Strong understanding of reliability, observability, security, and performance at scale.
Comfortable diving into complex technical problems and debugging production systems alongside your team.
A strong people leader who enjoys coaching engineers, providing feedback, and helping individuals grow.
Able to balance technical depth with effective delegation and team leadership.
Excited to work in a fast-moving startup where you will have significant ownership and ambiguity.
Bonus Points If You Have:
Experience managing infrastructure or platform teams supporting machine learning or LLM workloads.
Experience with infrastructure-as-code tools such as Terraform.
Experience with Kafka, Redis, PostgreSQL, ClickHouse, or similar distributed data systems.
Experience building internal developer platforms or platform engineering organizations.
Knowledge of cloud security, infrastructure hardening, or zero-trust architectures.
Experience scaling infrastructure through a period of rapid company or product growth.
Previous experience at a high-growth startup or building products from zero to one.
Interest in AI safety, cybersecurity, or adversarial machine learning.
If you don’t have 100% of these, you should still seriously consider applying. We care more about what you can do than your credentials.
You’ll Thrive Here If You:
You lead from the front. You're technically curious, comfortable getting into the details when needed, and know how to empower engineers without becoming a bottleneck.
You build great teams. You care deeply about developing engineers, creating a strong team culture, and giving people the ownership and autonomy they need to do their best work.
You think at scale. You enjoy designing resilient infrastructure, making thoughtful architectural decisions, and building systems that are secure, observable, and built to grow.
You balance strategy with execution. You can zoom out to define long-term technical direction while staying close enough to the work to help the team navigate difficult technical problems.
You collaborate across disciplines. You work effectively with machine learning engineers, security researchers, and product teams, understanding that great infrastructure enables everyone else to move faster.
You're excited by our mission and startup environment. You enjoy moving quickly, adapting to change, and helping build the infrastructure foundation for the future of secure AI.
What We Offer:
We offer a competitive compensation package designed to reward impact and incentivize growth. Our compensation philosophy is informed by our current valuation and recent industry data.
Compensation: $255,000 - $315,000 plus performance based bonus and meaningful equity package
Benefits:
401k with up to 4% matching
28 days annual leave (vacation + holidays)
Health, dental, and vision coverage
Catered lunches (Pittsburgh office)
Flexible work arrangements
Visa sponsorship available for exceptional candidates
Interview Process
🔎 Application review. We read everything; we’ll respond within 10 days.
🗣 Recruiter Screen. We learn about you; you learn about us.
🧑💻 Technical interview or Hiring Manager Interview.
💻 Role related assesment
🗣 Experience & culture interview
😇 Reference checks. We’ll reach out to 3-5 references that you provide.
📃 Offer. If it’s mutual, we move fast.
How to Apply
Submit your resume, link to your portfolio, and answer the questions on the application.