Member of Technical Staff, Inference Systems
- Salary
- $230K–$350KUSD
- Moves you to
- United States
- Support
- Visa sponsorship
- Posted
- Sep 25, 2026
About the Company
Our client is a seed-stage, stealth-mode startup building a high-performance AI inference platform from the ground up. The founding team brings deep AI-infrastructure experience and is tackling the hardest problems in LLM serving — scheduling, KV cache management, request routing, and the runtime systems that power model inference at scale — with Rust at the core of the stack. This is a ground-floor opportunity to shape the architecture of a fast, reliable inference system without legacy constraints.
Recently founded · Small founding team · Industry: AI infrastructure / LLM inference
The Role
You'll join a small, fast-moving team building a new inference system from scratch. This is a role for a systems engineer who lives and breathes inference internals — attention, KV cache, batching, scheduling — and wants to own the whole stack rather than a narrow slice.
What you'll be doing
- Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack.
- Design and implement KV cache management, prefix caching, and optimizations that cut latency and cost per token.
- Scale serving across GPUs and nodes, tackling multi-GPU and multi-node challenges directly.
- Profile, benchmark, and ship performance improvements across the entire inference pipeline.
- Work closely with the founding team on the core architectural decisions that define the platform.
Tech stack: Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL
Requirements
- 2–10 years of experience as a backend or distributed systems engineer.
- Hands-on experience building, operating, or optimizing LLM inference or serving systems at the engine, router, or runtime layer — beyond just calling hosted APIs.
- A deep working knowledge of transformer inference internals: attention, KV cache, batching, scheduling, and where the real bottlenecks are.
- Performance-critical backend or distributed systems experience where latency, throughput, and cost were first-order concerns.
- Hands-on time with a production inference engine (vLLM, SGLang, or TensorRT-LLM) plus strong systems-language skills (Rust, C++, Go, or systems-level Python/PyTorch). If you haven't used Rust yet, you should be able to become productive in it within a few weeks of joining.
Nice to Haves
- Time on an inference team at a model provider, accelerator vendor, or research lab, or open-source contributions to vLLM, SGLang, or Dynamo.
- A CS or systems degree from a strong program.
- Production Rust experience, CUDA/Triton kernel work, multi-GPU or multi-node serving (NCCL, NVLink, RDMA), prefix caching, speculative decoding, or prefill/decode disaggregation.
Why Join
- Own the architecture of a brand-new inference system with no legacy constraints.
- Solve the hardest problems in LLM serving, from KV cache management to request routing.
- A performance engineer's dream: latency and cost per token are the whole game.
- High visibility on a small, elite team where your code powers the core engine.
Details
- Location: Palo Alto, CA
- Work policy: Full-time, on-site five days a week
- Compensation: $230K–$350K + equity
- Visa sponsorship: Open to visa transfers (OPT, H-1B) and new sponsorships (new H-1B, TN)
- Employment type: Full-time