Davidjoseph Co logo

Member of Technical Staff, Inference Systems

Salary
$230K–$350K
USD
Moves you to
United States
Support
Visa sponsorship
Posted
Sep 25, 2026
Is this job info correct?

About the Company

Our client is a seed-stage, stealth-mode startup building a high-performance AI inference platform from the ground up. The founding team brings deep AI-infrastructure experience and is tackling the hardest problems in LLM serving — scheduling, KV cache management, request routing, and the runtime systems that power model inference at scale — with Rust at the core of the stack. This is a ground-floor opportunity to shape the architecture of a fast, reliable inference system without legacy constraints.

Recently founded · Small founding team · Industry: AI infrastructure / LLM inference

The Role

You'll join a small, fast-moving team building a new inference system from scratch. This is a role for a systems engineer who lives and breathes inference internals — attention, KV cache, batching, scheduling — and wants to own the whole stack rather than a narrow slice.

What you'll be doing

  • Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack.
  • Design and implement KV cache management, prefix caching, and optimizations that cut latency and cost per token.
  • Scale serving across GPUs and nodes, tackling multi-GPU and multi-node challenges directly.
  • Profile, benchmark, and ship performance improvements across the entire inference pipeline.
  • Work closely with the founding team on the core architectural decisions that define the platform.

Tech stack: Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL

Requirements

  • 2–10 years of experience as a backend or distributed systems engineer.
  • Hands-on experience building, operating, or optimizing LLM inference or serving systems at the engine, router, or runtime layer — beyond just calling hosted APIs.
  • A deep working knowledge of transformer inference internals: attention, KV cache, batching, scheduling, and where the real bottlenecks are.
  • Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-order concerns.
  • Hands-on time with a production inference engine (vLLM, SGLang, or TensorRT-LLM) plus strong systems-language skills (Rust, C++, Go, or systems-level Python/PyTorch). If you haven't used Rust yet, you should be able to become productive in it within a few weeks of joining.

Nice to Haves

  • Time on an inference team at a model provider, accelerator vendor, or research lab, or open-source contributions to vLLM, SGLang, or Dynamo.
  • A CS or systems degree from a strong program.
  • Production Rust experience, CUDA/Triton kernel work, multi-GPU or multi-node serving (NCCL, NVLink, RDMA), prefix caching, speculative decoding, or prefill/decode disaggregation.

Why Join

  • Own the architecture of a brand-new inference system with no legacy constraints.
  • Solve the hardest problems in LLM serving, from KV cache management to request routing.
  • A performance engineer's dream: latency and cost per token are the whole game.
  • High visibility on a small, elite team where your code powers the core engine.

Details

  • Location: Palo Alto, CA
  • Work policy: Full-time, on-site five days a week
  • Compensation: $230K–$350K + equity
  • Visa sponsorship: Open to visa transfers (OPT, H-1B) and new sponsorships (new H-1B, TN)
  • Employment type: Full-time

Similar jobs

Apply for this job