Senior DevOps with Python
- Hiring from
- Europe, Kyrgyzstan
- Work type
- Remote
- Posted
- Sep 29, 2026
Is this job info correct?
Join Akvelon — build products used by millions!
Akvelon is an IT company with 20+ years of experience and 1,200+ engineers across 15+ locations worldwide.
We work with both well-known global tech companies, including Microsoft, Facebook, Airbnb, Dropbox, and Pinterest, and with growing startups.
Our teams are involved in different types of engineering projects, from cloud solutions and AI/ML systems to big data, web, and mobile applications.
Since we are remote-first, our engineers work in distributed teams with flexible hours. We value ownership, clear communication, and the ability to take responsibility for your part of the work.
About the role
The project expands an advanced Kubernetes-based inference benchmarking framework to provide end-to-end performance measurements, cross-cloud comparisons, and updated model test scenarios.
It includes automated benchmarking for startup latency and scale-out behavior across multiple model families and hardware configurations, integrates dashboards and reporting, and supports next-generation autoscaling, scheduling, storage, and multi-host workloads to ensure reliable measurement of real-world inference performance.
Requirements
- Advanced Kubernetes expertise including deep understanding of pod lifecycle, deployments, services, autoscaling and troubleshooting, with hands-on experience in GKE
- Experience with Python, including automation, scripting and infrastructure skills
- Experience with observability, monitoring, logging, tracing and performance benchmarking
- Practical experience with GCP services including GKE, GCS and Cloud Monitoring / Logging
- MLOps experience including deploying and operating ML models, working with vLLM or similar frameworks and managing GPU workloads in Kubernetes
- Solid understanding of autoscaling with HPA, Metrics Server and basic knowledge of Cluster Autoscaler
- Experience designing and maintaining CI/CD pipelines using GitHub Actions
- Strong Python skills for automation, scripting and infrastructure tasks
- Knowledge of AWS and Azure cloud services
Responsibilities
- Provide consistent, platform-wide performance signals across all inference workloads and teams, ensuring clear visibility into system efficiency and bottlenecks
- Deliver standardized cross-cloud benchmarking across major Kubernetes providers to ensure reliable performance comparisons
- Support leadership reporting through monthly cloud-wide performance results that enable accurate insights and data-driven decision making
- Enable teams to validate new features including autoscaling, scheduling, storage, node provisioning, vLLM optimizations, and accelerator support using unified benchmarking frameworks aligned with organizational OKRs
- Extend the benchmarking framework to cover startup latency and scale-out behavior for multiple model families and hardware configurations
- Integrate automated benchmarking APIs, dashboards, and reporting pipelines to streamline performance evaluation
- Collaborate with engineering teams to maintain reusable inference components while ensuring accurate scheduling, infrastructure provisioning, and reporting
- Flexibility, ability to work with overlap until 11 AM PST