Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
MG

Inference Performance Engineer

Material Group
Posted 3 hours ago
🇺🇸United States🏢Hybrid📁Engineering & Development
Is this job info correct?

About the role Serving frontier models at scale requires solving novel systems problems at every layer of the stack. As an Inference Performance Engineer, you'll own the runtime that turns accelerators into a production serving system, optimizing throughput, latency, and cost across thousands of nodes. You'll work alongside hardware and compiler teams operating at the frontier of AI silicon design. What you'll do Build and improve the inference runtime Design scheduling, continuous batching, KV cache, and prefill/decode disaggregation Implement low-precision kernels and speculative decoding Drive throughput, latency, and cost per token Collaborate with hardware teams on kernels, operators, and graph optimizations Own the OpenAI-compatible API surface and serving protocol Build benchmarking, profiling, and regression infrastructure What you'll need BS in CS, EE, or related field, or equivalent experience Software engineering experience: Rust, Go, Python, or C++ Understanding of concurrency, memory, and tail latency Understanding of modern inference: transformers, attention, KV cache, batching, speculative decoding, quantization Experience with model serving frameworks: vLLM, TGI, SGLang, TensorRT-LLM, llama.cpp, or custom runtimes GPU or ASIC programming experience: CUDA, ROCm, Triton, or vendor-native toolchains Experience with low-precision inference (FP8, FP4, INT4) Profiling and benchmarking experience: Nsight, perf, custom harnesses What we offer Top-tier compensation structured to recognize and retain the best talent Meaningful equity Comprehensive medical, dental, vision, life, and disability insurance Parental leave for all new parents, including adoptive and surrogate journeys Flexible PTO Paid Holidays Relocation support Equal Employment Opportunity We're an Equal Opportunity Employer and do not discriminate on the basis of any protected status under applicable law.

Similar jobs

Similar jobs

Nvidia logo

Senior Systems Software Engineer - GPU Performance at Scale

Nvidia

🇺🇸United States10 hours ago
Nvidia logo

Senior Software Engineer, AI Performance Analysis

Nvidia

🇺🇸United States10 hours ago
Bright Vision Technologies logo

ML Performance Engineer

Bright Vision Technologies

🇺🇸United States13 hours ago
Pddn Inc. logo

Mainframe Performance Management Engineer on W2

Pddn Inc.

🇺🇸United States17 hours ago
Harvard University logo

Energy Performance Engineer

Harvard University

🇺🇸United States19 hours ago
Gevernova logo

Senior Engineer- Performance Engineering

Gevernova

🇺🇸United StatesYesterday