Designworks Talent logo

Applied Researcher – Network Expert

Designworks Talent
Posted 4 hours ago
United StatesHybridResearch & Science
Is this job info correct?

Applied Researcher – Network Expert

Location: Hybrid | Bellevue, WA (downtown)


About the Opportunity

Our client is seeking a Network Expert to join an Applied Research organization focused on the future of large-scale AI infrastructure.

This role is designed for a networking expert with deep experience in GPU-based computing environments and high-performance inter-GPU networking. As AI workloads increasingly rely on thousands of GPUs operating as a single logical system, the network connecting those GPUs becomes a critical component of overall infrastructure performance and scalability.

You will serve as a horizontal subject-matter expert, partnering with engineering and infrastructure leaders to evaluate technologies, architectures, and industry developments. You will help the organization understand where GPU networking is headed and what those developments mean for infrastructure, architecture, and investment decisions.

What You'll Do

  • Hold the company's view of where AI networking is going.

  • Track vendor and hyperscaler roadmaps, research, standards work, and the startup and venture landscape across scale-up, scale-out and scale-across fabrics, topology and optics, collectives and the software above the fabric, operations and telemetry, and tenant-facing capability. Right now that means questions like how fast Ultra Ethernet displaces the RoCEv2 fabrics most clusters actually run, where the scale-up domain should end now that NVLink has open challengers, what co-packaged optics does to power per port, and how to network a cluster that no longer fits in one building. Those specific questions will have changed within a year — holding the current version of them is the job.

  • Formulate and validate the product and engineering thesis. Turn that view into a defensible position on what we build, buy, or partner for in the fabric, pressure-tested against measured cluster performance, isolation requirements, and cost per port — and say so plainly when the evidence does not hold up.

  • Own the company's fabric position: where the scale-up and scale-out boundary sits for our workloads, which transport we bet on and when, and what good telemetry and observability look like so fabric problems are diagnosable rather than inferred. A fault that restarts a long training job is a direct cost, not an availability statistic.

  • Help finance to formulate the numbers: cost per port, optics and cabling economics across pluggables and co-packaged options — reach, power draw, failure rates at scale — and the performance we can actually substantiate against what vendor benchmarks claim.

  • Own tenant-facing network capability for GPU-as-a-service: multi-tenant isolation and its performance cost, storage traffic alongside GPU traffic, and what we can commit to contractually — including an honest read of where we are undifferentiated against peers.

  • Make the work land commercially. Support sales and delivery in demanding customer conversations about cluster performance, feed product and go-to-market with what we can offer at what performance and price, and provide technical diligence on network vendors, partners, and prospective tuck-in targets. Work closely with the Data Center Expert where the fabric meets the physical plant.


What We're Looking For

Required Qualifications

  • Deep professional experience in networking at scale,

  • Hands-on experience with high-performance GPU or HPC fabrics at current generations,

  • Demonstrated experience debugging real collective communication performance problems in production

Preferred Qualifications

  • Fluency with the landscape you would be scanning: the switch, NIC, and optics vendors, the standards bodies and consortia, and the startups attacking the fabric layer — and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.

  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources — vendor roadmaps, standards drafts, benchmark data, academic literature, your own testing — and producing a defensible position under genuine uncertainty. A PhD in a relevant technical field is one good route to this and is valued here; sustained industry research, standards-body work, or a body of internal technical assessments that changed real decisions are equally valid. Either way, the role turns on the second half: translating that work for engineering, product, go-to-market, and finance, because it informs all four.

  • Deep professional experience in computer networking and large-scale infrastructure.

  • Significant experience with GPU clusters, AI infrastructure, or high-performance computing environments.

  • Strong understanding of inter-GPU networking.

  • Hands-on experience with InfiniBand at current generations (NDR/XDR), high-performance Ethernet fabrics (RoCEv2, Spectrum-X, or Ultra Ethernet), or comparable HPC interconnects.

  • Practical understanding of NVLink / NVSwitch scale-up domains and how they interact with the scale-out fabric.

  • Experience debugging real collective communication performance problems — NCCL/RCCL, congestion, stragglers, topology mismatch — not just reading about them.

  • Understanding of multi-tenant network isolation and the security and performance tradeoffs involved in serving tenants on shared fabric.

  • Understanding of large-scale AI cluster configurations and the networking requirements associated with thousands of GPUs.

  • Ability to evaluate competing technologies and understand where the networking industry is heading.

  • Familiarity with NVIDIA and AMD GPU infrastructure ecosystems.

  • Strong analytical and technical communication skills.

  • Ability to operate as a horizontal technical expert and influence engineering decisions without necessarily owning implementation.

  • Ability to read technical papers and translate research concepts into practical engineering implications.

  • Ability to operate across engineering, research, and infrastructure organizations.

  • A track record of collaborating with researchers and engineers across groups and levels to shape long-term research directions and move research into practice.

  • Comfortable operating as an individual contributor with high ownership in a lean, early-stage team.

Location

  • Hybrid role based in downtown Bellevue, WA.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

  • Export control: this role involves technologies subject to U.S. export control regulations. Candidate eligibility may be subject to export control screening and, where applicable, licensing.

  • Travel: Willingness and ability to travel as needed internationally to data centers and co-locations (up to 25%)

Why Join?

  • High-impact technical role: Directly influence the technology direction of an organization building AI infrastructure at scale.

  • Ground-floor opportunity: Help establish technical strategy, architecture, processes, and culture within a growing organization.

  • High ownership: Operate as a senior individual contributor with substantial autonomy and direct access to senior technical leadership.

  • Cross-disciplinary exposure: Work across AI models, inference, accelerators, software systems, networking, infrastructure, and economics.

  • Cutting-edge technical problems: Work on multi-accelerator inference, intelligent routing, performance optimization, token economics, and compiler, kernel, and runtime technologies.

  • Research with practical impact: Turn emerging research and technology developments into decisions that directly affect engineering, product, commercial strategy, and investment.

  • Lean, senior environment: Work with a small group of highly experienced technical contributors rather than within a large management hierarchy.

Similar jobs