Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
LM Studio logo

Software Engineer, Inference Runtime

LM Studio
Posted 3 hours ago
🇺🇸United States🏢Hybrid💰$150.0K–$350.0K📁Engineering & Development
Is this job info correct?

LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family. As a team, we work with high technical intensity and personal responsibility. We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software. The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on. Qualifications Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure Strong programming ability in Python and C++ Deep understanding of transformer architectures and the mechanics of model inference Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution Takes personal responsibility for the correctness and performance of their work Bonus Qualifications Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM Responsibilities Maintain and push forward our inference stack on-device and in the cloud Bring up new model architectures and multimodal models Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution Benchmark and diagnose correctness and performance problems across the inference stack Contribute upstream to open-source projects such as llama.cpp and MLX Benefits Competitive salary and equity grants Great medical, vision, dental healthcare plans Catered team lunch / expensed dinners in the office Flexible PTO Flexible WFH Sun-drenched office in SoHo in NYC

Similar jobs

Similar jobs

Anthropic logo

Staff+ Software Engineer, Inference Runtime

Anthropic

🇺🇸United StatesJun 12, 2026, 8:49 PM UTC
Baseten logo

Software Engineer - Voice AI (Inference Runtime)

Baseten

🇺🇸United StatesMay 27, 2026, 7:35 PM UTC
Sonos logo

Senior Infrastructure Architect

Sonos

🇺🇸United States25 minutes ago
Sonos logo

Ecommerce QA Specialist

Sonos

🇺🇸United States25 minutes ago
IS

Detection and Response Lead

Integrated Specialty Coverages, LLC

🇺🇸United States1 hour ago
CLS-Group logo

Vice President, Detection Intelligence & Content Engineering

CLS-Group

🇺🇸United States1 hour ago