Iforce logo

Systems Engineer (Core Infrastructure)

Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 25, 2026
Is this job info correct?

Job Description

This is a remote position.

Role Overview

We are seeking an experienced Systems Engineer to design, deploy, and scale our high-performance cloud compute platform and distributed GPU infrastructure. In this role, you will be responsible for building and maintaining bare-metal hardware, high-density GPU clusters, virtualisation layers, and low-latency networking that power complex AI/ML workloads. You will work closely with Software, Architecture, and Product teams to drive the performance, reliability, and security of our customer-facing compute infrastructure.

Key Responsibilities

Compute Infrastructure & GPU Cluster Management

  • Provision, manage, and optimise large-scale bare-metal servers and specialised hardware configurations across multi-datacenter environments.

  • Architect and maintain high-performance compute clusters equipped with modern GPU accelerators, managing driver deployments, CUDA runtime environments, and firmware life-cycle operations.

  • Build and automate hypervisor environments (KVM, QEMU) and container orchestration platforms (Kubernetes) tailored for intensive parallel computing and multi-tenant isolation.

Network Engineering & High-Performance Storage

  • Implement and maintain ultra-low-latency network fabrics, Virtual Private Clouds (VPCs), and software-defined networking (SDN) solutions.

  • Deploy and scale high-throughput distributed storage architectures designed for data-heavy training and inference pipelines.

  • Optimise system performance across storage, memory, compute, and inter-node interconnects to ensure maximum platform utilisation.

Automation & Site Reliability

  • Develop infrastructure-as-code (IaC) configurations using tools such as Terraform, Ansible, or custom automation scripts to ensure rapid, reproducible bare-metal provisioning.

  • Design proactive monitoring, alerting, and observability frameworks to track system telemetry, hardware health, and thermal efficiency.

  • Participate in high-availability architecture planning, disaster recovery testing, and incident response to maintain stringent SLA targets.

Strategic & Technical Leadership

  • Evaluate emerging server hardware, accelerator technologies, and cloud orchestration frameworks to continuously refine system design.

  • Collaborate with Software and Platform Engineering teams to build seamless APIs and control planes on top of raw physical infrastructure.

  • Maintain comprehensive architecture documentation, hardware baseline specs, and operational runbooks for complex infrastructure workflows.

Requirements & Qualifications

Technical Expertise

  • Linux Kernel & Systems Administration: Deep, hands-on mastery of Linux systems (Debian/Ubuntu, RHEL/Rocky, Arch variants), kernel tuning, and low-level system performance optimisation.

  • Hardware & Accelerators: Proven experience managing enterprise server hardware, high-density server chassis, and GPU/accelerator acceleration platforms.

  • Virtualisation & Containers: Strong proficiency with Kubernetes, Docker, KVM, and cloud management frameworks (e.g., OpenStack, custom orchestrators).

  • Automation & IaC: Expertise in automated system provisioning, Configuration Management (Ansible, Puppet, or Chef), and Infrastructure-as-Code (Terraform).

  • Networking & Security: In-depth knowledge of BGP, VLANs, overlay networks, firewall policies, zero-trust architectures, and multi-tenant security isolation.

Mindset & Soft Skills

  • Strong analytical and troubleshooting skills under high-pressure, live operational conditions.

  • Passion for high-performance computing, open infrastructure, and scalable system design.

  • Clear communication skills with a proven ability to collaborate across software, architecture, and operational disciplines.



Similar jobs

Apply for this job