Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
Venice.ai logo

Engineering Lead, Inference Optimization

Venice.ai
Posted 7 hours ago
🇺🇸United States🏠Remote💰$270.0K–$330.0K📁Engineering & Development
Is this job info correct?

About Us Venice is the world’s leading consumer AI company built on principles of privacy, free speech, and user sovereignty. We’re building the Port City of AI, in which millions of individuals, third party apps, and AI agents gather, interact, and access sophisticated AI resources on a private and permissive foundation. Our mission is to make artificial intelligence approachable and useful in everyday work—bridging the gap between cutting-edge research and practical, real-world impact. We’re a fast-moving startup where every team member is expected to make a clear impact. Our culture is rooted in curiosity, ownership, ethical principle, philosophy, and collaboration—whether we’re designing better AI workflows, supporting our growing community, or shaping the future of human-AI interaction. Joining Venice AI means joining a team of unorthodox builders who believe in moving quickly, delivering a beautiful mass-marketand highly useful consumer product that doesn’t spy on people or censor their ideas and questions., and maintaining an edge in the rapidly evolving world of agentic machine intelligence. If you’re energized by big ideas, entrepreneurial spirit, individual empowerment, and the opportunity to help shape a fast-growing company in the world’s hottest industry from the ground up, please reach out.you’ll feel right at home here. Why we are hiring Venice is the only AI platform that runs inference with zero data retention and zero training on user inputs. This is an opportunity for you to be on the bleeding edge of privacy-focused AI with a unique and dedicated team of high-agency individuals alongside you. This role requires both hands on work as an individual contributor as well as the management of a small team. You will play a pivotal role, shaping Venice's overarching technical strategy and assembling an exceptional team to deliver peak inference performance at massive scale. The base annual salary for this position ranges from $270,000-$330,000 USD and reports to the Head of Engineering. What you'll do Own Venice’s technical strategy for inference performance Recruit and lead the Inference Optimization Team at Venice Optimize Venice's GPU infrastructure across a range of architectures (e.g. H200s, B300s) Improve latency, throughput, and cost per token for LLM inference workloads Build reproducible benchmarking harnesses across inference engines (e.g. vLLM, SGLang) to identify the optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU Work with our inference routing system to optimize multivariate inference load-balancing algorithms Evaluate emerging inference optimization techniques (custom CUDA/Triton kernels), novel attention variants, new quantization schemes, and compilation stack improvements. Hands-on kernel development experience is a strong plus. Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in Venice's stack. Who you are 8+ years in performance optimization or HPC, with deep GPU architecture and parallel programming knowledge 5+ years experience leading engineering teams Proficiency in Python, Rust, or Go. Bonus: C++/CUDA Hands-on experience with at least one production LLM inference engine (e.g. vLLM, SGLang) running at high volume in production Demonstrated experience with LLM inference optimization techniques: continuous batching, PagedAttention/KV cache management, speculative decoding, quantization, CUDA graphs, and torch.compile Fluency with quantization tradeoffs, both qualitative and quantitative Experience with distributed inference strategies (tensor parallelism, pipeline parallelism, MoE parallelism) in multi-GPU and multi-node environments Fluency with GPU profiling (Nsight Systems, Nsight Compute, PyTorch Profiler) and a bias toward measuring before optimizing Bonus: diffusion/image model inference optimization, custom Triton kernels, contributions to open-source inference frameworks

Similar jobs

Similar jobs

Dragonfly Careers logo

Senior Inference Optimization Engineer - Dragonfly Portfolio

Dragonfly Careers

🇺🇸United States6 days ago
DigitalOcean logo

Staff Engineer, Inference Optimizations

DigitalOcean

🇺🇸United States3 weeks ago
Modular logo

Inference Optimization Manager

Modular

🇺🇸United StatesJun 18, 2026, 7:58 AM UTC
Modular logo

Inference Optimization Engineer

Modular

🇺🇸United StatesJun 18, 2026, 7:58 AM UTC
Zoox logo

Senior AI Inference Engineer - Model Optimization & Deployment

Zoox

🇺🇸United StatesMay 27, 2026, 7:43 PM UTC
Cvshealth logo

Senior Software Development Engineer

Cvshealth

🌍Ireland, Israel, United States4 hours ago