Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education & Training jobs
  • Remote Healthcare & Nursing jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact mahmoud@relomote.com · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Netpreme logo

AI Infrastructure Engineer

Netpreme
Posted 7 hours ago
📦Relocation support🛂Visa sponsorship
🇺🇸United States
📁Data & Analytics
Is this job info correct?

About the Role We're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team. Essential Duties & Responsibilities Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines. Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems. Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization. Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding. Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation. Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime. Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability. Work with model/system engineers to bring newly released models into production efficiently. Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions. Contribute to defining next-generation benchmarks and service requirements as workloads evolve — multi-turn coding, agentic pipelines, RAG, and other long-context use cases. Qualifications BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience. 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering. Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang. Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution. Strong Python engineering skills. Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs. Clear written and verbal communication skills to work effectively with a small, fully distributed team. [ Preferred Qualifications (optional)] Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc. Experience with MoE / long-context model deployment. Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization. Experience with Nsight Systems / PyTorch Profiler. Familiarity with Kubernetes / production GPU serving. Previous startup experience. Compensation & Benefits Competitive salary with performance-based bonus and early-stage equity grant 100% employer-paid Health, Dental, and Vision coverage for you and your dependents 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing) Daily lunch stipend Enterprise-level Claude & ChatGPT access with a generous token budget Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office Visa sponsorship and relocation assistance to one of our office hubs A collaborative, continuous-learning environment with smart, dedicated colleagues building the next generation of high-performance computing architecture The Opportunity Impact: Humanity stands at the dawn of a new industrial revolution driven by AI—one with the potential to redefine how we live on this planet. We are tackling a fundamental challenge at the infrastructure layer: unlocking greater AI capability while dramatically improving efficiency. The work we do here compounds across state-of-the-art AI models, systems, and real-world applications. Timing: Breakthrough technology matters most when it meets the right time. Joining now means real ownership of the company and meaningful influence over product direction and execution. In this early-stage environment, your ideas shape the trajectory of the technology—not just its implementation. You’ll work from first principles, move quickly from insight to execution, and see your contributions directly reflected in what we build. Culture: You’ll work alongside a group of people who care deeply about rigor, clarity, and impact. We value thoughtful disagreement, fast learning, and intellectual fearlessness. This is a place where strong ideas shine, curiosity is encouraged, and growth is a daily practice—not a future promise.

Similar jobs

Similar jobs

Allen Institute logo

Senior Software Engineer - AI / ML Platforms and Infrastructure

Allen Institute

🇺🇸United StatesJul 21, 2026, 9:20 AM UTC
Davidjoseph Co logo

Causal Labs — Machine Learning Infrastructure Engineer

Davidjoseph Co

🇺🇸United States2 weeks ago
SK AX USA logo

Infrastructure Network Engineer

SK AX USA

🇺🇸United States2 weeks ago
SK AX USA logo

Cloud Infrastructure Engineer (Azure/AWS)

SK AX USA

🇺🇸United States2 weeks ago
Dedalus Labs, Inc. logo

Infrastructure Engineer Intern

Dedalus Labs, Inc.

🇺🇸United States3 weeks ago
Ivo Inc. logo

Infrastructure Engineer

Ivo Inc.

🇺🇸United States3 weeks ago