Client provides the world's best realtime AI models for consumer-scale applications. Its core product suite includes:
Who We're Looking For
we're looking for strong people from varied backgrounds who learn fast, thrive in ambiguity, and can show us what they've built, broken, and understood.
Experience We Find Useful
You don't need all of this. But you need enough to make a case.
Inference Optimization. Deep understanding of modern serving frameworks and techniques like vLLM or TRT-LLM.
Model Acceleration. Hands-on experience with quantization, distillation, caching strategies , continuous batching, paged attention, and speculative decoding.
High-Performance Systems. Proficiency in C++, CUDA, Rust, or highly optimized Python. Y
Distributed Systems & Scaling.
Experience with Kubernetes, Ray, custom load balancing, multi-GPU/multi-node inference, and reliably handling thousands of concurrent connections.
You can take a model from the research team, containerize it, optimize its serving, and ensure it runs reliably in production.
Background. PhD in CS, Physics, Math, or equivalent practical experience building backend or ML systems.
Tech stack
vLLM, TRT-LLM, CUDA, C++, Rust, Python, Kubernetes, Ray, NVIDIA GPUs, Docker
Solutions Architect - AI Inference Specialist
Friendliai
Associate/Journey Salesforce Developer, Information Technology (REMOTE AVAILABLE) [R0151805]
Nshe
Platform Engineer
Volarisgroup
QA Automation Engineer
Volarisgroup
Bswift System Analyst - Remote (EST)
Onedigital
Eligibility Technician Lead
Denver