Hosting Services Operations Engineer
- Hiring from
- United States
- Work type
- Hybrid
- Posted
Show job descriptionHide job description
Join Stanford University’s IT Hosting Services team as a Hosting Services Operations Engineer, supporting the on premise and cloud infrastructure that Stanford’s academic, research, and administrative communities depend on every day. This is a hybrid role: you will spend part of your time on organization-level cloud operations — helping manage cloud accounts, security guardrails, and cloud spend for hundreds of research and enterprise users — and part of your time hands-on in Stanford’s enterprise and research data centers, racking, stacking, cabling, and providing hands-and-feet support for the physical infrastructure behind those services. You will work closely with experienced engineers, security staff, and vendors building practical skills across public cloud, private cloud, and physical data center operations. It is an excellent role for someone who wants both depth in cloud platforms and real hardware experience in a large, mission-critical research environment.
Cloud Operations and Financial Operations (FinOps)
- Support multi-account cloud operations: Assist in the day-to-day operation of Stanford’s organization-level cloud footprint across AWS, Google Cloud, and Azure — including account and project provisioning, organizational unit and folder structure, and lifecycle tasks such as onboarding, transfers, and decommissioning.
- Apply and maintain guardrails: Help implement and monitor organization-wide policies — service control policies, organization policies, Azure Policy, tagging standards, and baseline configurations — so that new accounts land in a consistent, compliant state.
- Assist with landing zone operations: Support the standard account baseline (networking, logging, identity integration, budget alerts) used by enterprise customers, and help troubleshoot when accounts drift from it.
- Monitor cloud health and usage: Assist in monitoring service health, quotas and limits, and usage trends across the organization, escalating issues and helping coordinate responses with affected teams.
- Support identity and access: Help administer federated access, roles, and permission sets, applying least-privilege practices under the guidance of senior engineers.
- Assist with cloud automation: Participate in developing scripted and infrastructure-as-code workflows (for example Terraform, Python, or CLI-based tooling) that reduce manual effort in account management and reporting.
- Contribute to FinOps practice: Apply FinOps best practices under guidance to help Stanford departments and users get the most value per dollar from their cloud spend.
- Support cost visibility and allocation: Assist in maintaining tagging and account structures, cost allocation reports, and chargeback/showback data so that departments and administrators can clearly see where their cloud dollars go.
- Help identify savings opportunities: Assist with rightsizing analysis, idle and orphaned resource cleanup, storage class and lifecycle optimization, and workload scheduling recommendations for computing workloads.
- Support commitment and discount management: Help track and report on Reserved Instances, Savings Plans, Committed Use Discounts, and negotiated agreements, and assist in analyzing coverage and utilization.
- Assist with budgets, forecasting, and anomaly detection: Help configure budgets and alerts, investigate cost anomalies, and support forecasting for departmental workloads where predictable spend matters.
- Support user enablement: Help prepare cost reports, documentation, and guidance that make cloud economics understandable to users, staff, and students — including support for budgeting and credit programs.
Data Center Operations — Racking, Stacking, and Hands-and-Feet
- Install and decommission hardware: Rack, stack, cable, label, and decommission servers, storage arrays, and network equipment in Stanford’s enterprise and research data centers.
- Perform structured cabling work: Run and dress copper and fiber, manage patch panels and cross-connects, and maintain accurate port and circuit documentation.
- Provide hands-and-feet support: Serve as on-site hands for remote engineers, vendors, and research groups — swapping drives and optics, reseating components, power cycling equipment, console access, media handling, and smart-hands escorts.
- Support power and environmental needs: Assist with rack power planning and PDU connections (including DC power environments), monitor space, power, cooling, and weight constraints, and flag capacity concerns.
- Handle break/fix and RMAs: Diagnose basic hardware faults, open and track vendor support cases, coordinate parts replacement, and complete RMA returns.
- Maintain asset and facility records: Keep inventory, DCIM, and rack elevation records accurate, and support receiving, staging, and secure disposal of equipment.
- Participate in maintenance windows: Support scheduled maintenance, migrations, and after-hours or on-call work as needed.
Private Cloud and Systems Operations
- Support compute and storage infrastructure: Assist in monitoring the private cloud environment, helping to ensure system performance and availability.
- Maintain virtualization platforms: Support and troubleshoot virtualized environments, with guidance from senior engineers.
- Maintain system stability: Assist in monitoring system stability, responding to incidents, and ensuring minimal disruptions.
- Optimize resources: Assist in monitoring and analyzing resource usage, identifying potential improvements in compute and storage efficiency.
- Document processes: Help document infrastructure and data center procedures, runbooks, and diagrams for team reference and knowledge sharing.
You Will Also:
- Assist and contribute to the installation and maintenance of operating systems, utilities, and software on computing systems.
- Follow established protocols to assist in maintaining system and physical security.
- Collaborate with other IT teams, research computing groups, and facilities staff to confirm strategies, assess technical feasibility, and ensure compatibility within the university’s framework.
- Assist with capacity planning for system configurations, software services, network services, data center space and power, and load distribution.
Core Duties:
Build, install, configure, analyze, tune, and troubleshoot operating system to achieve optimum performance levels.
Resolve difficult system problems and provide consultation or training.
Configure and assist in the design of system and network security.
Manage hardware, software, and utilities for installation, modification, troubleshooting, maintenance, and upgrades of operating systems and workstation environments.
Monitor and analyze resource usage to recommend/develop enhancements to system capabilities and performance.
Compare, evaluate, and implement new technologies, and integrate systems into the computing environment.
Document systems infrastructure for users, support and consulting personnel, and developers.
Train personnel who provide support and consulting services to users.
Act as liaison with various departments across campus.
Facilitate vendor relationships.
Minimum Education and Experience:
Bachelor's degree and five years of relevant experience, or a combination of education and relevant experience.
Knowledge, Skills and Abilities:
Experience with multiple types of operating systems or expert knowledge in one.
Experience with managing an integrated computing environment.
Extensive experience in multiple programming languages for system development.
Experience with networked environments.
Experience with network storage solutions.
Experience with configuring and maintaining routers.
Demonstrated knowledge of security protocols.
Ability to manage shared resources and perform moderately complex tasks.
Ability to work well independently and as a team member.
Ability to train to others in applications and operating system fundamentals.
Experience with middleware and infrastructure.