Systems Engineer (Core Infrastructure)
- Hiring from
- Probably Worldwide
- Work type
- Remote
- Posted
- Sep 25, 2026
Job Description
This is a remote position.
Role Overview
We are seeking an experienced Systems Engineer to design, deploy, and scale our high-performance cloud compute platform and distributed GPU infrastructure. In this role, you will be responsible for building and maintaining bare-metal hardware, high-density GPU clusters, virtualisation layers, and low-latency networking that power complex AI/ML workloads. You will work closely with Software, Architecture, and Product teams to drive the performance, reliability, and security of our customer-facing compute infrastructure.
Key Responsibilities
Compute Infrastructure & GPU Cluster Management
-
Provision, manage, and optimise large-scale bare-metal servers and specialised hardware configurations across multi-datacenter environments.
-
Architect and maintain high-performance compute clusters equipped with modern GPU accelerators, managing driver deployments, CUDA runtime environments, and firmware life-cycle operations.
-
Build and automate hypervisor environments (KVM, QEMU) and container orchestration platforms (Kubernetes) tailored for intensive parallel computing and multi-tenant isolation.
Network Engineering & High-Performance Storage
-
Implement and maintain ultra-low-latency network fabrics, Virtual Private Clouds (VPCs), and software-defined networking (SDN) solutions.
-
Deploy and scale high-throughput distributed storage architectures designed for data-heavy training and inference pipelines.
-
Optimise system performance across storage, memory, compute, and inter-node interconnects to ensure maximum platform utilisation.
Automation & Site Reliability
-
Develop infrastructure-as-code (IaC) configurations using tools such as Terraform, Ansible, or custom automation scripts to ensure rapid, reproducible bare-metal provisioning.
-
Design proactive monitoring, alerting, and observability frameworks to track system telemetry, hardware health, and thermal efficiency.
-
Participate in high-availability architecture planning, disaster recovery testing, and incident response to maintain stringent SLA targets.
Strategic & Technical Leadership
-
Evaluate emerging server hardware, accelerator technologies, and cloud orchestration frameworks to continuously refine system design.
-
Collaborate with Software and Platform Engineering teams to build seamless APIs and control planes on top of raw physical infrastructure.
-
Maintain comprehensive architecture documentation, hardware baseline specs, and operational runbooks for complex infrastructure workflows.
Requirements & Qualifications
Technical Expertise
-
Linux Kernel & Systems Administration: Deep, hands-on mastery of Linux systems (Debian/Ubuntu, RHEL/Rocky, Arch variants), kernel tuning, and low-level system performance optimisation.
-
Hardware & Accelerators: Proven experience managing enterprise server hardware, high-density server chassis, and GPU/accelerator acceleration platforms.
-
Virtualisation & Containers: Strong proficiency with Kubernetes, Docker, KVM, and cloud management frameworks (e.g., OpenStack, custom orchestrators).
-
Automation & IaC: Expertise in automated system provisioning, Configuration Management (Ansible, Puppet, or Chef), and Infrastructure-as-Code (Terraform).
-
Networking & Security: In-depth knowledge of BGP, VLANs, overlay networks, firewall policies, zero-trust architectures, and multi-tenant security isolation.
Mindset & Soft Skills
-
Strong analytical and troubleshooting skills under high-pressure, live operational conditions.
-
Passion for high-performance computing, open infrastructure, and scalable system design.
-
Clear communication skills with a proven ability to collaborate across software, architecture, and operational disciplines.