Please Note: This role is based at the Stanford Historic Campus, Pine Hall. The Engineering Manager, Automation & Service Reliability leads a team of four engineers who design, build, and support the integration platforms, robotic process automation (RPA), and data services for Stanford University's IT Infrastructure team. This is a hybrid technical working manager and people leadership role. The manager actively contributes to technical projects, sets technical direction, and reviews engineering work while also leading the team's day-to-day people management, including coaching, performance management, hiring, and career development. The manager partners closely with the Director of Communications Technologies Service Support and represents the team with internal University IT stakeholders and healthcare partners including Stanford Health Care (SHC) and Stanford Medicine Children's Health (SCH). This role also carries a forward-looking mandate: evolve the team's automation practice beyond traditional RPA and integration work toward agentic AI. The manager will identify, pilot, and scale agentic AI solutions — AI agents and LLM-powered workflows that can reason, plan, and take multi-step action across existing systems — to solve problems for clients across the broader IT Infrastructure organization, not just the team's existing service lines. Key Responsibilities People Management Lead, manage and grow a team of one Technical Lead and three Software Engineers including training, goal-setting, performance reviews, and individual development plans. Recruit, hire, onboard, and mentor engineering talent; build a culture of technical excellence and accountability. • Balance workload across the team, manage on-call/support coverage, and resolve interpersonal or performance issues. Technical Leadership Guide architecture and design decisions across AI solutions, Mulesoft integrations, UiPath automations, Oracle APEX applications, and data pipelines. Review code and designs at a level sufficient to ensure reliability, security, and maintainability, without necessarily being the primary hands-on developer. Own service reliability for production systems (SIPdb, TDC, data feeds) including incident response, root-cause analysis, and preventive engineering. Establish and maintain an intake process for new automation requests, prioritizing the team's backlog against ServiceNow requests and strategic initiatives. Agentic AI & Innovation Define and drive a roadmap for applying agentic AI — AI agents and LLM-powered workflows capable of multi-step reasoning and action — across the team's existing platforms (Mulesoft, UiPath, Oracle APEX) and newly identified use cases. Identify high-value opportunities for agentic AI across the IT Infrastructure organization, working with peer teams to surface pain points suited to AI-driven automation. Stand up pilots for agentic AI use cases (e.g., ticket triage and resolution, incident summarization, service desk drafting, data feed and reporting automation) and define success metrics before scaling to production. Evaluate AI tooling, platforms, and integration patterns (including LLM APIs and MCP-style tool/agent frameworks), and build organizational guardrails around security, data handling, and responsible use. Upskill the team in AI-assisted and agentic development practices, and champion the team's evolution toward an AI-forward automation practice (e.g., an Automation Center of Excellence model). Stakeholder & Program Collaboration Serve as the team's primary technical point of contact for IT Infrastructure, University Campus and Healthcare partners. Collaborate with the Director and broader Communications Technologies leadership on change management for next-generation voice and contact center initiatives. Communicate roadmap, risk, and status clearly to both technical and non-technical audiences. Core Duties : Lead and manage all business, technical, and education activities (e.g. application development, installing, configuring, and maintaining servers, routers, firewalls, workstations, and network equipment). Exercise full management responsibility for a technical group, including recruiting, hiring, training, developing, evaluating, and setting priorities. May manage managerial staff. Ensure work completion within schedule, budgetary, and design constraints; make decisions about analysis, design, and testing; solve complex technical problems; provide alternative methods for achieving goals when necessary. Provide strategic planning for own work group; may assist higher level management in broader scope strategic planning for a large, complex, university-wide function or major initiative. Create procedures and guidelines to ensure compliance with university policy and federal and state regulations. Develop and manage budgets for projects or work groups. May oversee or assist in preparation and submission of documentation, such as proposals, progress reports, or other contractual requirements. Monitor technology trends and evaluate emerging technologies for adoption and implementation. Work collaboratively with colleagues to leverage the university/school’s investments in information technology. Minimum Education and Experience: Bachelor's degree and five years of relevant work experience, or a combination of education and relevant experience. Knowledge, Skills and Abilities : Detailed understanding of relevant technical knowledge and problem resolution. Strong customer relationship skills, consensus building skills, and ability to establish effective working relationships in a diverse environment. Demonstrated ability to lead, motivate, and develop staff.
Service Reliability Engineer
Nvidia
Service Reliability Engineer
Nvidia
Enterprise Service Reliability and Insights Lead
Decisionpointcorp
Manager, Site Reliability Engineering - Storage Layer Service
MongoDB
Senior Site Reliability Engineer, Messaging Services
National Basketball Association (NBA)
Site Reliability Engineer (Senior or Staff), Storage Layer Services (SLS)
MongoDB