AI Platform Engineer The AI Platform Engineer is responsible for deploying and managing the computing and storage platform backbone that powers AI models. Working on-site alongside AI Data Engineering, IT Infrastructure, Security, and business stakeholders, this role is central to the reliability, performance, and scalability of our AI/ML infrastructure. Responsibilities: Provision and maintain CPU and GPU compute clusters, high-bandwidth network equipment, and scalable, high-throughput storage systems for AI/ML training and inference, working on-site and virtually to coordinate hands-on configuration and troubleshooting with infrastructure teams. Develop and maintain scripts and infrastructure-as-code (IaC) modules for automated configuration of AI/ML infrastructure using tools such as Terraform and Ansible; participate in in-person and virtual code and design reviews with the platform engineering team. Develop and maintain CI/CD pipelines and container orchestration platforms (Kubernetes, Flux); troubleshoot pipeline failures and escalate to more senior engineers or other teams as needed, collaborating in person and virtually to resolve issues quickly and effectively. Set up and monitor data stores (file systems, block storage, traditional and vector databases) to support LLMs and RAG pipelines, with regular on-site coordination with AI Data Engineering teams to ensure data infrastructure meets evolving model requirements. Monitor system performance and log metrics to ensure reliability and uptime; implement system improvements under guidance of more senior engineers, including attending in-person and virtual planning and incident review sessions. Collaborate in person with AI Data Engineering, IT Infrastructure, Security, and Business Leaders to optimize infrastructure for performance and cost, including participating in cross-functional working sessions, architecture reviews, and stakeholder briefings. Deploy and maintain Model Context Protocol (MCP) connectors and ensure secure, scalable infrastructure for MCP integrations, coordinating directly with security and engineering teams to validate configurations and address emerging requirements. Basic Qualifications: At least two (2) years of experience working in an infrastructure, DevOps, or platform engineering role, with at least some exposure to AI/ML workloads. Proficiency in cloud computing, container orchestration, and infrastructure-as-code. Knowledge of networking and distributed storage; Experience with monitoring and observability tools to track performance and system health. Ability to troubleshoot complex systems and collaborate in person with Data Engineering, Infrastructure, AI Engineering, Security, and Business Leaders. Strong problem-solving, communication, and teamwork skills; comfortable working in a collaborative, on-site team environment. Foundational understanding of AI/ML concepts. Preferred Qualifications: Knowledge of GPU acceleration. Travel: 10%
Senior Software Engineer - Research Platform, Consumer Devices
OpenAI
Managing Engineer , Database & Platform - Remote
Allstate
Microsoft Azure Cloud Platform Engineer Expert
Allstate
Staff Software Engineer - AI Platform Engineer
Wellsky
Lead Cloud & Data Platform Engineer
Humana
Senior Platform Engineer
Collibra