Client provides the world's best realtime AI models for consumer-scale applications. Its core product suite includes:
Who We're Looking For
we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood.
Experience We Find Useful
You don't need all of this. But you need enough to make a case.
Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
Model Acceleration. Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding.
High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. Y
Distributed Systems & Scaling.
Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections.
You can take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production.
Background. PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.
Tech stack
vLLM, TRT-LLM, CUDA, C++, Rust, Python, Kubernetes, Ray, NVIDIA GPUs, Docker
Forward Deployed Engineer, Simulations [33026]
Stealth Startup
Software Engineer
Jack
Software Engineer, AI/ML [33018]
Stealth Startup
Founding Member of Technical Staff - Platform Engineering
Halluminate
Founding Member of Technical Staff - Research / Post-Training
Halluminate
Electronic Systems Verification & Validation Project Team Leader
Cat