Infrastructure Engineer (GPUaaS – AI Neocloud)
- Hiring from
- Australia
- Work type
- Remote
- Posted
Show job descriptionHide job description
🗓️ Full-time | Permanent | Start ASAP
📍 Australia wide | Sydney preferred
🏡 Remote-first (ANZ) | Hybrid option in Sydney
About Sharon AI
Sharon AI is an Australian neocloud, delivering trusted AI infrastructure organisations need to build, train and run AI at scale. We support customers across the full AI lifecycle, from training through to inference and agentic AI, drawing on a strong ecosystem of technology and co-location partners to deliver capability where it's needed. As the first neocloud to join NVIDIA's AI Compute Program, we're growing quickly, scaling our AI Factory platform to meet rising demand for advanced compute.
The Role
As Sharon AI continues to grow, we're looking for an Infrastructure Engineer (GPUaaS – AI Neocloud) to join our Infrastructure Engineering team and help build and deploy the core infrastructure powering our AI Factories and AI Clouds, reporting to our Infrastructure Engineering Team Lead.
In this hands-on role, you'll build scalable, reliable and efficient infrastructure for GPU-intensive AI/ML workloads, taking new AI Factory and AI Cloud capacity from initial build through to production readiness. You'll work closely with senior infrastructure engineers, platform engineers, observability engineers and project managers, with exposure to cutting-edge technology at the forefront of AI infrastructure.
What you'll be doing
· Build and deploy multi-tenant infrastructure supporting GPU-as-a-Service (GPUaaS) platforms.
· Build and deploy high-performance compute clusters using Kubernetes and/or HPC schedulers such as Slurm.
· Manage GPU node lifecycle activities including provisioning, configuration, patching and decommissioning.
· Develop and maintain Infrastructure-as-Code using Terraform, Ansible or similar tooling, and automate routine provisioning, scaling and configuration tasks.
· Monitor GPU utilisation and scheduling, helping identify opportunities to improve efficiency, performance and cost.
· Maintain and extend monitoring, logging and alerting, while supporting high availability, backup and disaster recovery processes.
· Apply security controls and workload isolation for multi-tenant environments, and work with observability and platform teams to validate new deployments.
· Provide L3 support for platform and infrastructure incidents, contribute to root cause analysis, and maintain accurate documentation, runbooks and operational procedures.
· Contribute to capacity planning and ongoing infrastructure improvements while learning from and sharing knowledge with the wider engineering team.
What we're looking for
· 3–5 years’ experience in infrastructure engineering, systems engineering, platform engineering or SRE roles.
· Strong Linux systems administration and troubleshooting skills, with solid knowledge of cloud and/or bare-metal infrastructure environments.
· Hands-on experience with Kubernetes, containers and production infrastructure systems.
· Working knowledge of Infrastructure-as-Code such as Terraform or Ansible, plus automation, CI/CD and Git.
· Programming or scripting skills in Python, Go or Bash, with an understanding of networking and storage fundamentals.
· Familiarity with observability tools and practices such as Prometheus and Grafana.
· A security-aware, proactive and collaborative approach, with strong communication skills and the ability to work effectively in a fast-paced environment.
· A genuine interest in GPU infrastructure, AI/ML technology and what Sharon AI is building.
· Bachelor’s degree in computer science, Engineering or a related field, or equivalent practical experience.
Nice to have:
· Experience with GPUaaS, IaaS or neocloud platforms, or exposure to GPU-based systems and workloads.
· Familiarity with AI/ML workloads and frameworks such as PyTorch or TensorFlow, HPC schedulers such as Slurm or Ray, and GPU technologies including CUDA, NCCL or MIG.
· Exposure to high-performance networking such as RDMA or InfiniBand, and/or relevant cloud, Linux or Kubernetes certifications.
Why Join Sharon AI?
You'll be joining a rapidly growing Australian technology business at an exciting stage of its journey, with the opportunity to work directly with the infrastructure and technology powering the next generation of AI.
🏡 Hybrid working – flexibility between our office and working from home
🎂 Birthday leave – take some extra time to celebrate your day
🧠 Employee Assistance Program (EAP) – confidential support when you need it
🤖 Exposure to AI and next-generation technology – work in one of the fastest-moving areas of technology
🎓 Learning & development – we support your career growth with approved conferences, professional memberships & courses
🚗 Novated leasing – a tax-effective way to finance and run your car
🎁 Bounty referral program – generous rewards for successfully referring new talent to Sharon AI
🏆 Employee of the month – recognition plus a $500 gift card
🌏 Growing global business – be part of an Australian technology company with an expanding international footprint
💡 Make an impact – join at a stage where your work can have a visible influence on how we build and grow
🤝 Collaborative culture – work alongside talented people across technical, commercial and corporate teams
Our Values
Integrity | Innovation | Collaboration | Wellbeing | Inclusion
Apply today and help us build the infrastructure powering the next generation of AI.
Due to the high volume of applications we receive, we’re unfortunately not always able to provide individual feedback to unsuccessful candidates. We appreciate your understanding and want to assure you that every application will be reviewed with care.