Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Inferact logo

Member of Technical Staff, TPU Performance Engineering

Inferact
Posted Jun 26, 2026, 9:15 AM UTC
🛂Visa sponsorship
🇸🇬Singapore
📁Engineering & Development
Is this job info correct?

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build. About the Role We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware. You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable. Skills and Qualifications Minimum qualifications: Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar. Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling. Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads. Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths. Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work. Preferred qualifications: Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems. Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems. Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation. Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs. Bonus points if you have: Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure. Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads. Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements. Logistics Location: This role is based in Singapore. Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity. Visa sponsorship: We sponsor visas on a case-by-case basis. Benefits : Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Similar jobs

Similar jobs

BAH Partners logo

C++ Engineers | High-Performance AI Infrastructure | Equity Options | Singapore

BAH Partners

🇸🇬Singapore2 days ago
Aaru logo

Solutions Engineer, APAC

Aaru

🇸🇬Singapore2 days ago
MiAO AI logo

Founding Engineer, AI Agents

MiAO AI

🇸🇬Singapore3 days ago
MiAO AI logo

Software Engineer, AI Agents

MiAO AI

🇸🇬Singapore3 days ago
HU

Forward Deployed Research Engineer

hud

🌍United States, Singapore1 weeks ago
HU

Research Engineer, Benchmarks

hud

🌍United States, Singapore1 weeks ago