EPAM Systems logo

Senior HPC DevOps Engineer

Moves you to
Poland
Support
Relocation support
Posted
Is this job info correct?
Show job description

We are seeking a Senior HPC DevOps Engineer to join our Science, Innovation & Labs team, responsible for the scaling, reliability, and automation of our high-performance computing (HPC) and machine learning operations (MLOps) platform.

Responsibilities

  • Guide scientists and data teams in navigating and utilizing the platform user interface (UI) effectively, helping them run self-service workloads without direct infrastructure friction
  • Advise users and manage infrastructure capacity regarding capacity blocks versus on-demand usage, optimizing cost, quotas, and resource availability for heavy workloads
  • Maintain automated pipelines for infrastructure provisioning and platform service deployments
  • Resolve technical queries regarding job scheduling failures, cluster bottlenecks, and resource quotas
  • Collaborate with developer experience teams to improve documentation
  • Collaborate with engineering teams to monitor GPU utilization via tools such as CloudWatch or Prometheus
  • Manage AWS GPU instance families and allocate block compute for large-scale ML training and inference pipelines
  • Ensure compute availability through capacity planning and reservation management
  • Deploy containerized environments tuned for HPC and GPU pass-through
  • Deploy and scale HPC workloads on cloud infrastructure utilizing parallel storage and networking solutions

Requirements

  • 5+ years of experience in HPC or DevOps engineering roles
  • Knowledge of MPI, OpenMP, and multi-node GPU communication protocols such as NCCL and GPUDirect
  • Proven experience managing AWS GPU instance families, including P-series, G-series, and Tranium/Inferentia
  • Hands-on mastery of AWS Capacity Blocks for ML, On-Demand Capacity Reservations (ODCRs), and Service Quota management
  • Experience in deployment of containerized environments using Apptainer/Singularity, Docker, or Enroot
  • Understanding of I/O performance bottlenecks when interfacing with distributed file systems such as Lustre, GPFS, BeeGFS, or AWS FSx for Lustre
  • Hands-on skill in profiling applications using NVIDIA Nsight or similar tools to locate memory and compute bottlenecks
  • Experience deploying or scaling HPC workloads on cloud infrastructure utilizing EFA, ParallelCluster, and parallel storage (FSx for Lustre)
  • Proficiency in English at a B2+ level

We offer

We gather like-minded people:

  • Top tech minds driving innovation in AI, cloud and digital platform modernization
  • Supportive team and agile, startup-like culture
  • Hybrid by design mode and opportunity to work remotely within Poland
  • Chance to work abroad for up to 60 days annually
  • Business-driven relocation opportunities

We provide growth opportunities:

  • Career development programs
  • Thought leadership, mentoring, soft skills and well-being programs
  • Certification (Anthropic, Gemini, GCP, Azure, AWS)
  • English classes

We cover it all:

  • Stable pay
  • Participation in the Employee Stock Purchase Plan with a 15% discount
  • Benefits package (health insurance, multisport, shopping vouchers)
  • Referral bonuses up to $2,000
  • Offices featuring entertainment and relaxation zones, table tennis and football, free snacks, coffee and more
  • Corporate, social and well-being events

Please, note:

  • Benefits listed above are available to employees only
  • We are open for working with Contractors. Terms of B2B cooperation agreements are agreed individually
  • We will reach out to selected candidates exclusively

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI-Native enterprises, driving measurable value from innovation and digital investments.

Remote in Poland

Similar jobs

Apply for this job