AI Platform Engineer L3 (GPUaaS – AI Neocloud)
📍 EMEA | Remote-first
About Sharon AI
Sharon AI is building the infrastructure powering the next generation of artificial intelligence.
Operating across AI infrastructure, high-performance compute, cloud platforms and large-scale technology environments, Sharon AI delivers scalable, secure and reliable infrastructure for demanding AI, ML and HPC workloads.
The Role
As an AI Platform Engineer L3, you'll design, build and operate the platform layer powering Sharon AI's GPU-as-a-Service (GPUaaS) offering across the EMEA region. You'll own the architecture, automation and reliability of the platform services sitting above Sharon AI's GPU and network fabric, spanning Kubernetes, Slurm, GPU scheduling, MLOps tooling, model serving and platform observability across multiple EMEA sites.
Reporting to the Head of Operations, you'll work closely with Network Engineering, Infrastructure and customer-facing teams to solve complex, cross-team challenges and ensure Sharon AI's platform can scale reliably and efficiently across the region. This is a hands-on senior individual contributor role suited to an engineer with strong platform engineering, DevOps or MLOps experience who can operate independently in a fast-paced, AI-native environment and provide technical guidance to less experienced engineers.
Key Responsibilities
- Design and own the AI platform architecture across Sharon AI's EMEA GPU clusters, including Kubernetes, Slurm and container orchestration
- Lead the development of CI/CD pipelines and MLOps tooling supporting training, fine-tuning and inference workloads across multiple sites
- Define and implement multi-tenant GPU resource scheduling, quota management and workload isolation strategies at scale
- Own the design of model serving infrastructure, balancing high availability, performance and cost efficiency
- Build and evolve platform-wide observability across monitoring, logging and alerting, covering platform health, GPU utilisation and workload performance
- Drive Infrastructure-as-Code adoption and platform automation using Terraform and Ansible
- Partner closely with Network Engineering to integrate the platform layer with Sharon AI's underlying InfiniBand/RDMA fabric
- Act as a senior escalation point for complex platform issues impacting customer AI/ML workloads across EMEA
- Lead incident response and post-incident reviews, contributing to operational runbooks and platform best practice
- Partner with the Head of Operations on platform capacity planning, scaling strategy and cost optimisation across EMEA
- Mentor and provide technical guidance to less experienced platform engineers
- Support enterprise GPUaaS customer onboarding and technical escalations across the region
Skills & Experience
- 6–10+ years' experience in platform engineering, DevOps, MLOps or SRE, ideally within HPC, cloud or AI/ML infrastructure environments
- Bachelor's degree in Computer Science, Electrical Engineering or a related field
- Hands-on experience operating Kubernetes and GPU scheduling at production scale
- Proven experience designing CI/CD and Infrastructure-as-Code practices for platform teams
- Proven experience supporting GPU or AI/ML infrastructure at scale, ideally within a GPUaaS or neocloud environment
- Deep expertise in Kubernetes and GPU scheduling frameworks, including Slurm, Kubernetes device plugins and NVIDIA GPU Operator
- Strong experience designing and operating MLOps tooling and ML pipeline orchestration in production
- Advanced proficiency in Python, Bash and Infrastructure-as-Code tools such as Terraform and Ansible
- Strong understanding of GPU infrastructure and distributed training concepts, including NCCL, data/model parallelism and RDMA-aware scheduling
- Experience with platform-wide observability tooling such as Prometheus, Grafana and telemetry stacks
- Strong Linux systems and networking fundamentals, with the judgement to independently solve ambiguous, cross-team problems
- A security-first mindset when operating within multi-tenant environments
- Awareness of EMEA regulatory and data residency considerations, including GDPR, as they relate to platform operations
- Strong communication and collaboration skills across distributed, multi-region teams, with the ability to mentor other engineers
Experience with distributed training frameworks such as PyTorch and TensorFlow, MLOps platforms including MLflow, Kubeflow or Ray, and InfiniBand/RDMA or RoCEv2 networking concepts is advantageous. Kubernetes certifications such as CKA, CKAD or CKS, and cloud certifications across AWS, GCP or Azure, are also advantageous. The role requires the right to work in an EMEA jurisdiction, with existing eligibility to work across the EU/EEA or UK advantageous given the multi-country remit.
Why Join Sharon AI
- Own the platform architecture powering a growing GPU-as-a-Service and AI neocloud business across EMEA
- Work hands-on with large-scale GPU infrastructure, Kubernetes, Slurm and AI-native platform technologies
- Shape how Sharon AI's platform scales across multiple jurisdictions, sites and customer workloads
- Solve complex technical challenges across GPU infrastructure, MLOps, networking and distributed AI workloads
- Influence platform reliability, automation, capacity and cost optimisation across the region
- Work closely with Network Engineering and Infrastructure teams on high-performance InfiniBand/RDMA environments
- Provide technical leadership and mentorship while remaining hands-on as a senior individual contributor
- Help enterprise customers reliably train, fine-tune and run AI/ML workloads at scale
- Join a highly technical and ambitious team operating at the forefront of AI infrastructure
Our Values
Integrity | Innovation | Collaboration | Wellbeing | Inclusion