Software Engineer ML infrastructure
- Salary
- €70K–€120KEUR
- Hiring from
- Europe
- Work type
- Remote
- Posted
- Sep 29, 2026
At Orient Path, we partner with high-growth technology companies globally.
One of our clients is an early-stage deep-tech AI company developing a fundamentally new approach to LLM model compression — making foundation models smaller, faster, and more efficient to run in production.
Their research team develops compression methods. Now they are building the engineering layer that turns this research into reliable, high-performance ML infrastructure.
What are we building?
Production infrastructure around LLM compression and inference.
The engineering work sits between models and hardware: quantization, pruning and retraining pipelines, inference frameworks, GPU execution, performance profiling, benchmarking, and production tooling.
Who are we looking for?
A Software Engineer — ML Systems & AI Infrastructure who has already built and run the systems behind modern ML/LLM workloads.
This is not a general backend, MLOps, or AI application role.
You should be strong in at least one of these areas:
- Training: distributed model training, FSDP/ZeRO/DDP, mixed precision, checkpointing and multi-GPU workloads
- ML Infrastructure: LLM serving, GPU clusters, distributed execution, profiling, throughput and production ML systems
- GPU Performance / Kernels: GPU profiling, memory and compute optimization, kernel performance, CUDA/Triton or similar low-level work
Across all three, we expect strong software engineering fundamentals and hands-on ownership of real ML systems.
Tech stack:
Python, PyTorch, vLLM, Hugging Face, TensorRT-LLM, llama.cpp, CUDA/Triton, GPU profiling & benchmarking, distributed GPU systems, CI/CD
What we expect from candidates:
- Strong production Python and software engineering fundamentals
- Hands-on experience building or optimizing ML/LLM systems, not simply using AI APIs or models
- Strong understanding of GPU performance: memory, compute, profiling, bottlenecks and hardware constraints
- Experience with LLM inference, distributed training, GPU infrastructure, or kernel optimization
- Experience with PyTorch and modern ML/LLM frameworks
- Ability to investigate performance problems hands-on: profile → identify the bottleneck → implement the change → measure the result
- Experience turning research/experimental code into reusable, reliable engineering systems
- Understanding of quantization, pruning, mixed precision, model compression or other model-efficiency techniques
Especially relevant experience:
CUDA/Triton or GPU kernels; vLLM/TensorRT-LLM/llama.cpp; FSDP/ZeRO/DDP; NCCL and multi-GPU systems; profiling with Nsight or similar tools; open-source ML infrastructure contributions.
What we offer:
- €70K–120K base salary, depending on experience
- Equity participation
- Vienna / remote within Europe
- Visa sponsorship & relocation support
- Direct work with founders and researchers
- High ownership in a small, deeply technical team
If this sounds close to what you've actually been building, I'd love to talk.