Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
EER Poland logo

Inference Stack Engineer

EER Poland
Posted Jun 3, 2026, 5:07 PM UTC
🇵🇱Poland🏢Hybrid📁Engineering & Development
Is this job info correct?

Inference Stack Engineer Join us as an Inference Stack Engineer to shape the runtime between models and hardware: design, optimize, and own high-performance AI execution at scale. Apply here! Inference Stack Engineer (AI Systems / Compiler & Runtime) We are building a next-generation AI inference stack designed for high-performance execution on modern and custom compute architectures. Our mission is to deliver industry-leading low-latency and high-throughput AI systems by designing and optimizing the full execution path — from model representation to hardware-level execution. This is a deeply technical role at the intersection of compiler systems, AI runtimes, and high-performance computing . You will work on core infrastructure that defines how modern AI models are executed efficiently at scale. What you will do Design and build components of an AI inference stack , from high-level model representation to low-level execution Develop and extend a Python-based DSL for expressing AI workloads and kernels Work on compiler infrastructure including: IR design and transformation pipelines graph lowering and optimization passes backend code generation for target execution environments Optimize model execution for: latency throughput memory efficiency numerical stability Contribute to runtime systems responsible for model execution and scheduling Profile and analyze inference workloads to identify system bottlenecks Collaborate closely with hardware and systems engineers on execution efficiency Influence architecture decisions for next-generation AI execution platforms What we are looking for Strong software engineering background (C++ and Python) Experience with performance-critical systems or compiler-related work Understanding of AI model execution (especially transformers / LLMs) Familiarity with compute graphs, tensor operations, or execution frameworks Ability to analyze complex systems end-to-end (model → runtime → hardware) Experience working with large codebases and system-level debugging Strong communication skills and ability to work in cross-functional teams Nice to have Experience with compiler frameworks such as: LLVM MLIR Triton TVM XLA Experience contributing to deep learning frameworks (PyTorch, TensorFlow, JAX) Understanding of GPU or accelerator execution models Experience with kernel optimization or operator-level performance tuning Knowledge of distributed inference systems (e.g. NCCL, RPC-based serving) Familiarity with hardware-aware optimizations (memory hierarchy, vectorization, scheduling) What we offer Work on the core execution layer of modern AI systems Direct impact on inference performance of large-scale AI workloads Collaboration with experts in compilers, systems, and AI infrastructure Highly technical environment with strong engineering autonomy Opportunity to shape the architecture of a next-generation inference stack Competitive compensation and flexible working model Why this role is different This is not a typical ML engineering or application role. You will not be training models. You will be working on how models actually run efficiently , at scale, across compute systems, shaping the performance layer that sits between AI models and hardware. Locations Gdańsk Remote status Hybrid Employment type Full-time

Similar jobs

Similar jobs

Sensor Tower logo

Full Stack Engineer

Sensor Tower

🌍Canada, Poland, Portugal, Turkey, United Kingdom22 hours ago
Simcorp logo

Full-Stack Software Engineer

Simcorp

🇵🇱PolandYesterday
Simcorp logo

Lead Full-Stack Software Engineer

Simcorp

🇵🇱PolandYesterday
NC

Senior Full Stack Engineer

Netwrix Corporation

🇵🇱PolandYesterday
CL

Senior Fullstack Engineer (Angular+Node.js)

Codest Ltd. Company No. 12590542, VAT number: GB363431020

🇵🇱PolandYesterday
ExergyIQ logo

Senior Full-Stack Engineer (Java + React)

ExergyIQ

🇵🇱PolandYesterday