Observability Engineer
Sharon AI, Inc๐๏ธ Full-time | Permanent | Start ASAP
๐ Sydney preferred, Australia-wide remote considered
๐ก Hybrid working (if Sydney based)
About Sharon AI
Sharon AI is an Australian neocloud, delivering trusted AI infrastructure organisations need to build, train and run AI at scale. We support customers across the full AI lifecycle, from training through to inference and agentic AI, drawing on a strong ecosystem of technology and co-location partners to deliver capability where it's needed. As the first neocloud to join NVIDIA's AI Compute Program, we're growing quickly, scaling our AI Factory platform to meet rising demand for advanced compute.
The Role
As Sharon AI continues to grow, we're looking for a talented Observability Engineer to join our Infrastructure team and take ownership of end-to-end observability across our high-performance computing (HPC), AI infrastructure and data centre environments, reporting to our Head of Operations.
In this role, you'll design and build the monitoring, telemetry and automation capabilities that give us visibility into compute, storage, networking and facility infrastructure, enabling proactive performance optimisation, rapid incident resolution and continuous service improvement. You'll work closely with Infrastructure Engineering, Data Centre Operations, Service Management and Customer Success teams, combining strong technical expertise with analytical rigour to drive operational excellence and data-driven decision-making across the organisation.
What You'll Be Doing
- Design and maintain observability frameworks across compute, storage, networking and physical infrastructure environments
- Deploy and manage telemetry platforms for metrics, logs and traces using industry-standard observability tools
- Develop dashboards, visualisations and reporting capabilities that provide actionable operational insights
- Monitor and analyse HPC and AI workloads, including GPU, CPU, storage and network performance, to identify bottlenecks and optimisation opportunities
- Support performance benchmarking, capacity planning and infrastructure scaling initiatives
- Develop intelligent alerting and anomaly detection capabilities, support incident response and conduct root cause analysis for critical events
- Build automation solutions for telemetry collection, analysis and operational reporting, and integrate observability into CI/CD and infrastructure-as-code practices
- Partner with engineering, operations, service management and customer-facing teams to translate complex technical data into clear, actionable recommendations
- Maintain documentation, standards and operational procedures, supporting ITIL-aligned Incident, Problem and Change Management processes
What We're Looking For
- Bachelor's degree in Computer Science, Information Technology, Engineering or a related technical discipline
- 5โ10 years' experience in Observability Engineering, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Operations or related roles
- Proven experience supporting large-scale distributed computing, cloud, HPC or enterprise infrastructure environments
- Strong understanding of HPC architectures, including GPU and CPU compute environments, cluster-based workloads and parallel computing frameworks
- Experience with observability platforms such as Prometheus, Grafana, OpenTelemetry, Splunk, ELK Stack or equivalent
- Experience monitoring high-performance networking technologies, including InfiniBand, RDMA and low-latency networking environments
- Strong scripting or programming capability using Python, Bash, Go or similar technologies
- Working knowledge of ITIL service management principles, particularly Incident, Problem and Change Management
- Strong stakeholder management and communication skills, with the ability to engage effectively across technical and non-technical audiences
Nice to have:
- Experience operating within hyperscale, enterprise data centre, cloud infrastructure or AI computing environments
- ITIL v4 Foundation certification
- Certifications relating to observability, cloud platforms, Kubernetes or infrastructure technologies
- Experience with HPC workload schedulers such as Slurm
- Experience with Kubernetes, container orchestration and cloud-native infrastructure platforms
- Familiarity with DCIM platforms, environmental monitoring systems and facility telemetry integration
- Knowledge of AI infrastructure environments, GPU clusters and large-scale machine learning workloads
Why Join Sharon AI?
You'll be joining a rapidly growing Australian technology business at an exciting stage of its journey, with the opportunity to work directly with the infrastructure and technology powering the next generation of AI.
๐ก Hybrid working โ flexibility between our office and working from home
๐ Birthday leave โ take some extra time to celebrate your day
๐ง Employee Assistance Program (EAP) โ confidential support when you need it
๐ค Exposure to AI and next-generation technology โ work in one of the fastest-moving areas of technology
๐ Learning & development โ we support your career growth with approved conferences, professional memberships & courses
๐ Novated leasing โ a tax-effective way to finance and run your car
๐ Bounty referral program โ generous rewards for successfully referring new talent to Sharon AI
๐ Employee of the month โ recognition plus a $500 gift card
๐ Growing global business โ be part of an Australian technology company with an expanding international footprint
Our Values
Integrity | Innovation | Collaboration | Wellbeing | Inclusion
Apply today and help us build the infrastructure powering the next generation of AI.
Due to the high volume of applications we receive, we're unfortunately not always able to provide individual feedback to unsuccessful candidates. We appreciate your understanding and want to assure you that every application will be reviewed with care.