Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Featherlessai logo

Machine Learning Engineer — Inference Optimization

Featherlessai
Posted May 27, 2026, 8:45 PM UTC
🌍Worldwide🏠Remote📁Engineering & Development
Is this job info correct?

About the Role We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale . You’ll work at the intersection of research and production—turning cutting-edge models into fast, reliable, and cost-efficient systems that serve real users. This role is ideal for someone who enjoys deep technical work, profiling systems down to the kernel/GPU level, and translating research ideas into production-grade performance gains. What You’ll Do Optimize inference latency, throughput, and cost for large-scale ML models in production Profile and bottleneck GPU/CPU inference pipelines (memory, kernels, batching, IO) Implement and tune techniques such as: Quantization (fp16, bf16, int8, fp8) KV-cache optimization & reuse Speculative decoding, batching, and streaming Model pruning or architectural simplifications for inference Collaborate with research engineers to productionize new model architectures Build and maintain inference-serving systems (e.g. Triton, custom runtimes, or bespoke stacks) Benchmark performance across hardware (NVIDIA / AMD GPUs, CPUs) and cloud setups Improve system reliability, observability, and cost efficiency under real workloads What We’re Looking For Strong experience in ML inference optimization or high-performance ML systems Solid understanding of deep learning internals (attention, memory layout, compute graphs) Hands-on experience with PyTorch (or similar) and model deployment Familiarity with GPU performance tuning (CUDA, ROCm, Triton, or kernel-level optimizations) Experience scaling inference for real users (not just research benchmarks) Comfortable working in fast-moving startup environments with ownership and ambiguity Nice to Have Experience with LLM or long-context model inference Knowledge of inference frameworks (TensorRT, ONNX Runtime, vLLM, Triton) Experience optimizing across different hardware vendors Open-source contributions in ML systems or inference tooling Background in distributed systems or low-latency services Why Join Us Real ownership over performance-critical systems Direct impact on product reliability and unit economics Close collaboration with research, infra, and product Competitive compensation + meaningful equity at Series A A team that cares about engineering quality, not hype

Similar jobs

Similar jobs

Featherlessai logo

AI Researcher — Inference Optimization

Featherlessai

🌍WorldwideMay 27, 2026, 8:45 PM UTC
BR

Cyber Software Engineer

Breakpoint Research

🌍Worldwide5 hours ago
SM

Digital PR & Media Outreach Expert (Freelance)

Space M Online

🌍Worldwide1 hour ago
YM

Senior CRO Strategist

Yamu Media

🌍Worldwide1 hour ago
YM

Email Marketing Strategist

Yamu Media

🌍Worldwide1 hour ago
Matrix_Professional_Staffing_Solutions_Inc. logo

Legal Administrative Assistant

Matrix_Professional_Staffing_Solutions_Inc.

🌍Worldwide5 hours ago