Social Discovery Ventures logo

Lead ML Engineer - LLM Inference & NLP

Salary
$1K
USD per year
Hiring from
Serbia
Work type
Remote
Posted
Sep 28, 2026
Is this job info correct?

Social Discovery Group (SDG) is a group of social discovery companies. SDG solves the problems of loneliness, isolation, and disconnection - transforming virtual intimacy into the new normal. SDG’s products redefine the way people interact and connect with one another.

Our portfolio includes social entertainment platforms designed to connect people online across different cultures and regions of the world.

We bring together a team of like-minded people and IT professionals who specialize in creating and developing globally impactful social discovery products. Our international team of digital nomads works remotely from all over the world.

We’re proud to be a two-time “Great Place to Work” winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).

We are looking for Lead ML Engineer -LLM Inference & NLP.

Your main tasks will be:

  • Speed up and scale LLM inference in production: SGLang, KV and prefix caching, batching, quantization, speculative decoding
  • Run distributed inference for very large models (up to 1T+ parameters) across multi-GPU and multi-node setups
  • Benchmark new GPU servers and hardware, bring them into production and adapt our serving code to them
  • Lead the NLP and CV teams technically: review experiments, set direction, and step in early when something is heading the wrong way
  • Train and fine-tune the language models, and improve the agent harnesses and chat algorithm that run on them
  • Track cutting-edge research and open-source work in inference and post-training, and turn it into the ML roadmap
  • Collaborate closely with the validation, content, and dataset preparation teams to design experiments and measure model quality

    We expect from you:

    • Deep hands-on experience optimizing LLM inference in production with SGLang, vLLM, or TensorRT-LLM
    • Experience with distributed inference or training of large models: MoE, tensor/expert/pipeline parallelism, multi-node GPU clusters
    • Strong understanding of what makes inference fast: KV cache, attention kernels, batching, quantization, GPU profiling
    • Experience training and fine-tuning LLMs, including post-training (RLHF, DPO, or similar)
    • Proven technical leadership: you've guided engineers through reviews, mentoring and technical decisions while still writing code yourself
    • Proficiency with PyTorch, transformers, and related libraries
    • Experience at AI-focused startups or companies (Character AI, OpenAI, and similar is a strong plus)
    • Backend engineering experience (Python, Go, C#) and knowledge of scalable deployment systems is a significant advantage
    • Advanced English or Russian

    Nice to have:

    • CUDA or Triton kernel development
    • A computer vision background. We also welcome strong CV leads who have accelerated large generative image or video models
    • Experience with multimodal LLMs
    • First-author papers or notable open-source work, e.g. contributions to SGLang, vLLM, or post-training libraries
    • A degree in CS, math, or physics from a strong program (MSc or PhD)


    What do we offer:

    • REMOTE OPPORTUNITY to work full-time;
    • The initial pay level or pay range for this role will be shared with candidates during the recruitment process and before the commencement of employment;
    • Vacation 28 calendar days per year;
    • 7 wellness days per year (time off) that can be used to deal with household issues, to lie down and recover without taking sick leave;
    • Bonuses up to $5000 for recommending successful applicants for positions in the company;
    • 50% payment for professional training, international conferences, and meetings;
    • Corporate discount for English lessons;
    • ​Health benefits. According to the paychecks, if you are not eligible for corporate medical insurance, the company will compensate you with up to $ 1,000 gross per year per employee. This can be spent on self-purchase of health insurance or on doctor’s fees for yourself and close relatives (spouse, children);
    • ​Workplace organization. The company provides all employees with an equipped workplace and all the necessary equipment (table, armchair, wifi, etc.) in our offices or co-working locations. In the other locations, the company provides reimbursement of workplace costs up to $ 1000 gross once every 3 years, according to the paychecks. This money can be spent on the rent of the co-working room, on equipping the working place at home (desk, chair, Internet, etc.) during those 3 years;
    • Internal gamified gratitude system: receive bonuses from colleagues and exchange them for our merchandise, team building activities, massage certificates, etc.

    Sounds good? Join us now!

    Similar jobs

    Apply for this job