About The Role Volta builds and operates large scale GPU compute infrastructure for AI workloads. A single cluster is tens of thousands of GPUs, hundreds of switches, and tens of thousands of cables. At that size you cannot manage the fabric by hand, and you cannot find out whether a design works by deploying it. This role builds the model and the tooling that make large topologies tractable: a machine readable source of truth for every site, generated configuration that flows from it, and a simulated fabric where changes are proven before they touch hardware. You will write the software that lets a handful of engineers run fabrics that would otherwise need a room full of people. What You Will Be Doing Own Volta's network source of truth: the data model covering sites, racks, devices, interfaces, cabling, addressing, and rail-optimized topology, and keep it authoritative rather than descriptive. Generate device configuration from that model for every platform in the estate, so that config is an output of the model and never edited in place. Build and operate a simulation environment that reproduces full cluster topologies, and make pre-deployment validation of fabric change a normal step rather than an exception. Build the CI pipelines that validate network change: schema and policy checks, generated config diffs, simulated convergence, and reachability and routing assertions before merge. Detect and close drift between intended state and device state across sites, and make divergence visible rather than discovered during an incident. Automate bring-up verification with the bring-up teams: cable plan generation, LLDP based cabling validation, link quality and error checks, and acceptance test suites that produce a pass or fail against the design. Build the tooling that turns a new site from a design document into a provisioned fabric, and shorten how long that takes with each deployment. Instrument the fabric: streaming telemetry collection, topology aware metrics, and tooling that lets the team reason about a fabric of this size. Write production Python or Go in shared repositories, under the same review, testing, and CI standards as the rest of platform engineering. Support incident response and root cause work where modeling, simulation, or config history helps explain what happened. What You Bring 4+ years in network automation, infrastructure software, or network engineering with a substantial software component. Strong Python in production: testing, packaging, code review, and CI. Our working languages are Python, Go, and Rust. Data modeling experience with a network source of truth such as NetBox or Nautobot, including extending the model rather than only consuming it. Configuration as code in practice: templated or programmatic generation, declarative and idempotent workflows, and version controlled change. Network fundamentals at depth: L2/L3, VLANs, BGP, ECMP, leaf spine design, and overlay protocols. You need to understand what you are modeling. Hands-on experience with network simulation or emulation, for example containerlab, vendor virtual appliances, or an equivalent lab automation approach. Comfort operating at scale: thousands of endpoints and hundreds of devices, where anything that does not generate or validate automatically does not happen. Clear written communication. The model and its tooling are used by people who did not build them. Nice to Have (But Not Essential) None of these are required. Strength in one or more helps. Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery. Configuration analysis or formal verification tooling such as Batfish. NVIDIA Air, SONiC virtual switch, or Cumulus based lab environments. gNMI, OpenConfig, or NETCONF/YANG for configuration and telemetry. GPU fabric exposure: RoCE v2, InfiniBand, UFM and its API, or NCCL level performance validation. Go or Rust, or interest in moving further toward systems level languages. Graph based topology modeling, or experience where topology is queried rather than diagrammed. Kubernetes operators or controller patterns, and integrating network state into a control plane. Ansible, Nornir, or similar frameworks, with a view on where they stop being the right tool. Open source contributions to networking or automation projects.
QA Automation Engineer
Volarisgroup
Senior Security Automation Engineer
Fntg
Junior Security Automation Engineer
Fntg
Principal QA Automation Engineer
Morningstar
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Nvidia
QA Engineer – Automation
Bright Vision Technologies