RE

Senior Network Production Engineer

Hiring from
United Kingdom
Work type
Remote
Posted
Is this job info correct?

537,429 remote jobs, straight from company career pages

100% free · New jobs every hour

Show job description

Senior Production Engineer (Network Development)

Hyperscale AI Infrastructure (UK, Remote)

Up to £170K + Equity


We're hiring Production Engineers (x3) to own the health of a rapidly expanding AI data centre network. The company is an extremely well-funded AI infrastructure provider operating multiple hyperscale sites, with new capacity coming online continuously and a multi-gigawatt roadmap. This is a small founding UK team, working remotely as part of a global engineering organisation of over 300 people.


At this scale, the network can't be debugged by hand. You'll own fleet health end to end, and most of what you'll build doesn't exist yet.


Monitoring. Design the realtime monitoring, alerting and health dashboards for the whole network, giving on-call engineers an accurate view of network state across every site.


Debugging tooling. Build diagnostics for switch-to-switch and host-to-switch links, fleet-wide remote execution, and visual tooling that shows what's broken and why. The tooling has to separate optics failures from routing misconfiguration and power events, so manual triage stops being the default.


Repair automation. Build an automated pipeline from fault detection through vendor RMA, ticketing and optics inventory tracking, back to service, across fabric, edge and host networking.


The stack

  • Go, with Python for automation
  • gNMI, gRPC, NETCONF and SONiC for telemetry and device management
  • A strong plus: large-scale fabric experience (BGP, ECMP, spine-leaf) and out-of-band management


Who does well here

  • Network engineers who build software, or software engineers who know networks deeply
  • Network Development Engineers, Network Reliability Engineers or Production Engineers from hyperscalers or large network operators
  • People who've built telemetry pipelines, config generation, zero-touch provisioning or fault validation tooling, rather than just used them
  • Optical engineers who write code and want to move closer to the DC fabric
  • Python-first automators who haven't used Go yet: how fast you learn matters more than an exact match to the stack


How it works

  • Shared on-call: you lead incidents and fix root causes, not just symptoms
  • A fast pace, real ownership and very little red tape
  • A short, fast interview process, done remotely, with no take-home exercises


If you've built tooling that made network faults faster to find and fix, get in touch for the full details.


Similar jobs

Apply on LinkedIn