Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
VO

Senior Network Engineer

Volta
Posted 3 hours ago
🇺🇸United States🏢Hybrid💰$200.0K–$240.0K📁Engineering & Development
Is this job info correct?

About The Role Volta builds and operates large scale GPU compute infrastructure for AI workloads. The network is not a layer underneath our platform, it is part of it. Fabric design, overlay and multi-tenancy, edge connectivity, and the software that programs and observes all of it sit in one platform engineering team, deliberately. This is a network role with a software expectation attached. You will need real depth in fabric and routing, because a 20,000 GPU training cluster punishes anyone who understands the network only from documentation. You will also need to write production code, because at this scale anything that requires a human at a CLI does not happen reliably. Configuration is generated from a model, validated in CI, and applied by tooling. If your current work is mostly logging into devices and making changes by hand, this role will be a significant shift. You will work alongside platform engineers, not adjacent to them: same repositories, same review standards, same definition of done. What You Will Be Doing Common across the team: Design, build, and operate the network fabric across our sites: leaf-spine Ethernet, BGP underlay, EVPN and VXLAN overlay, and multi-tenant isolation. Build and extend the software that manages the fabric: configuration generation from a source of truth, validation pipelines, drift detection, and the tooling that makes change safe at scale. Contribute production Python or Go to the platform codebase, including the integration between the physical fabric and the IaaS control plane. Own how network capability is exposed upward: the APIs and abstractions through which tenants receive isolated, performant networking. Instrument the fabric: streaming telemetry, topology-aware metrics, and tooling that makes a fabric of this size understandable. Work with the bring-up teams during cluster deployment: fabric build, validation, acceptance testing, and turning the pain points you find into platform features rather than tribal knowledge. Debug the hard problems: congestion and packet loss under collective communication load, and the class of failure where controller state and hardware state disagree. Take part in on-call, incident response, and the follow-up work that closes structural gaps rather than only the immediate issue. Evaluate designs and configurations proposed by OEMs and partners, and challenge them where they do not fit our requirements. Participate in code review, technical design discussions, and cross-team collaboration in an Agile (Kanban or Scrum) environment. Depending on your background, you will go deeper in one of these areas: GPU fabric: RoCE v2 and InfiniBand for training traffic, congestion control tuning, rail-optimized topology, and the performance validation that proves a cluster is fit for workloads. SDN and overlay integration: the software boundary between the fabric and the platform, controller and API-driven fabric programming, and multi-tenant network provisioning. Edge connectivity: external connectivity, BGP peering, transit and IX relationships, and our public autonomous system. Fabric observability: telemetry pipelines, fabric health tooling, and making network state queryable rather than inspectable. What You Bring 4+ years in data center or cloud network engineering, in production environments where downtime had real consequences. Ethernet fabric depth: leaf-spine design, BGP including unnumbered BGP, ECMP, and the day-two realities of operating it. EVPN and VXLAN in production, including what happens to an overlay under multi-tenant load. Production software development, not scripting. Python or Go in a shared repository, under normal review, testing, and CI standards. Tools only you can run are not what we mean. Network automation in practice: configuration as code, a source of truth system such as NetBox or Nautobot, and declarative or idempotent workflows. Solid Linux fundamentals and comfort at the command line as an engineer, not only as an operator. Familiarity with Kubernetes networking and how workload networking interacts with the underlying fabric. Multi-vendor capability. Able to work across major OEM platforms and not dependent on one vendor's CLI. Willing to be on site during cluster bring-up when it matters. Clear written communication. Designs, decisions, and failure analysis need to be readable by people who were not in the room. Nice to Have (But Not Essential) None of these are required. Several map to specific areas of the team's scope, so strength in one or more helps us place you well: Fluency with AI-assisted development: agentic CLI tools, IDE assistants, and orchestrating multiple coding agents through MCP, skills, or APIs to amplify delivery. RoCE v2 at scale: PFC and ECN tuning, DCQCN, and how it behaves differently from InfiniBand under training load. InfiniBand production experience: fat tree topology, UFM, fabric partitioning, adaptive routing, and SHARP. NVIDIA Spectrum-X, including NetQ and Cumulus, or SONiC and whitebox platforms. gNMI, OpenConfig, or NETCONF and YANG for configuration and telemetry. IPv6 at production scale: dual stack design, v6 BGP peering, and addressing architecture. ASN operations: public autonomous systems, transit and IX peering, RPKI and IRR hygiene, and DDoS posture. OVN and OVS, SR-IOV, DPDK, or BlueField DPU based networking. Network simulation or emulation with containerlab, NVIDIA Air, or equivalent. Familiarity with NVLink and NVSwitch topologies and NCCL behavior. Depth in Go or Rust beyond working proficiency. Open source contributions to networking or infrastructure projects. Experience working distributed across time zones with counterparts in other regions.

Similar jobs

Similar jobs

ON

Sr. IT Project & Network Engineer

Oneoncology

🇺🇸United States3 hours ago
Northwell logo

Senior Network Engineer (Senior Enterprise Communications Engineer )

Northwell

🇺🇸United States3 hours ago
CT

Network Engineer

Core Technology Solutions

🇺🇸United States3 hours ago
VO

Network Modeling / Automation Engineer

Volta

🇺🇸United States3 hours ago
EC

Principal Network Engineer

Eci

🇺🇸United States3 hours ago
EC

Network Implementation - Senior Network Engineer, Implementation

Eci

🇺🇸United States3 hours ago