Relomote
Remote JobsRelocation Jobs
Add companySaved
Relomote

Relomote is a job board for remote, hybrid, and relocation jobs — every listing AI-classified for the countries it actually hires from, or the visa and relocation support it offers.

LinkedInCrunchbase

Remote jobs by category

  • Remote Engineering & Development jobs
  • Remote Customer Support jobs
  • Remote Design jobs
  • Remote Marketing jobs
  • Remote Sales jobs
  • Remote Product jobs
  • Remote Data & Analytics jobs
  • Remote People & Talent jobs
  • Remote Writing & Content Creation jobs
  • Remote Finance jobs
  • Remote Legal & Compliance jobs
  • Remote Operations & Admin jobs
  • Remote Data Entry jobs
  • Remote Virtual Assistant jobs
  • Remote Education/Training jobs
  • Remote Healthcare/Clinical jobs
  • Remote Other jobs

Remote jobs by location

  • Work from anywhere jobs
  • Remote jobs in Africa
  • Remote jobs in Asia
  • Remote jobs in Europe
  • Remote jobs in Latin America
  • Remote jobs in Middle East
  • Remote jobs in North America
  • Remote jobs in Oceania
  • All remote jobs →

Relocation & visa sponsorship

  • Visa sponsorship jobs
  • Relocation package jobs
  • Relocate to Europe
  • Relocate to Germany
  • Relocate to Netherlands
  • Relocate to Spain
  • Relocate to Portugal
  • Relocate to Greece
  • Relocate to United Kingdom
  • Relocate to Canada
  • Relocate to Australia
  • Relocate to Sweden
  • Relocate to Switzerland
  • Relocate to Japan
  • Relocate to United Arab Emirates
  • All relocation jobs →

© 2026 RelomoteAboutPrivacyTerms

Contact [email protected] · Built by Mahmoud

Relomote
Remote JobsRelocation Jobs
Add companySaved
NA

Senior Software Engineer - Distributed Systems

NeuroSpark AI
Posted 2 hours ago
🛂Visa sponsorship
🇺🇸United States
💰$170.0K–$350.0K📁
Engineering & Development
Is this job info correct?

NeuroSpark builds and operates a high-performance AI inference platform that helps enterprises run large language models faster, cheaper, and at scale. Inference infrastructure is the foundation the entire AI application layer runs on — every AI product ultimately depends on how fast, how reliably, and how affordably models can serve their users. Our vision is to make that layer so efficient that compute is never the reason a good AI product fails.


About the Role

Serving inference at scale is a scheduling problem. Requests arrive with wildly different shapes and latency expectations, GPUs are heterogeneous and expensive, and the difference between a platform that's fast and one that's economical usually comes down to how well work gets placed. That system is what you'll own.


You'll design and build the scheduling and routing layer of our platform: how requests get admitted, prioritized, batched, and placed across a heterogeneous multi-cloud GPU fleet, under real multi-tenant load and real latency commitments. This is core-systems work with a clean slate — you'll be making the foundational architectural decisions, not maintaining someone else's, and the quality of those decisions will show up directly in our margins and our customers' latency numbers.


You'll work close to the metal and close to the math. Some days that means reasoning about queueing behavior and control loops on a whiteboard; other days it means profiling Go or Rust until the tail latency comes down. We're a small team, so you'll own systems end-to-end — design, implementation, rollout, and the production reality afterward.


Responsibilities

  • Own the scheduling and routing layer — design and build request admission, prioritization, batching, and placement across a heterogeneous GPU fleet spanning multiple clouds and accelerator types
  • Engineer for latency and utilization at once — drive down tail latency while driving up fleet utilization; these fight each other, and resolving that tension well is the job
  • Model the system, not just code it — apply queueing theory, control theory, and load-shedding principles to make the platform behave predictably under bursty, multi-tenant traffic
  • Build multi-tenant fairness and isolation — ensure priority guarantees and SLO commitments hold when the fleet is saturated and customers are competing for the same capacity
  • Own it in production — instrument, observe, and debug distributed behavior in a live system; carry your designs through rollout and real-world load
  • Set the technical bar — make foundational architecture decisions, write the design docs that anchor them, and raise the engineering standard of everyone around you


Qualifications

This is a senior individual-contributor role. We're looking for someone who has built systems like this before and can operate independently from the first week.

  • Substantial experience building and operating large-scale distributed systems in production — you've owned something load-bearing, not just contributed to it
  • Track record of designing core systems from zero to one, and living with the consequences of your architectural decisions
  • Hands-on experience with scheduling, load balancing, request routing, or resource allocation systems
  • Strong systems fundamentals — operating systems, networking, concurrency — and the ability to reason quantitatively about system behavior using queueing theory, control theory, or similar
  • Fluency in a performance-sensitive language (Go, Rust, or C++), with the profiling and optimization instincts that come from actually chasing latency in production
  • Comfort with GPU infrastructure and LLM inference fundamentals — batching, KV cache behavior, throughput/latency tradeoffs; deep expertise here is a plus, but strong distributed-systems judgment matters more
  • Clear technical writing — you can make a hard design decision legible to people who weren't in your head
  • An AI-native way of working — you use AI tools daily and have your own view of how they change how infrastructure gets built


Preferred Qualifications

  • Mandarin proficiency is a plus


Nice to have: Kubernetes and multi-cloud operations experience; open-source contributions to inference, serving, or scheduling projects; experience operating GPU clusters at scale.


Why This Role

  • Ownership of a core system at the foundation of the platform, with the architectural latitude that comes with building it first
  • Meaningful equity — we expect the people who build the foundational systems to own a real piece of what they build
  • A small, high-caliber team where the distance between a good idea and it running in production is measured in days

Compensation and Benefits

The base salary range for this position is $170,000 – $350,000 per year. The range reflects the position across experience levels; actual base salary will be determined by job-related knowledge, skills, experience, and work location, and may fall anywhere within the stated range.


In addition to base salary, this position is eligible for equity in the company, along with medical, dental, and vision coverage, and other benefits.


Additional Information

We are able to sponsor H-1B and other work visas for qualified candidates.

Similar jobs

Similar jobs

Clera logo

Senior Software Engineer, Distributed Data Systems

Clera

🇺🇸United States2 days ago
DL

Distributed Systems Engineer

Dedalus Labs

🇺🇸United States1 weeks ago
Geico logo

Sr Staff Software Engineer - (Platform Engineering/Java/Distributed Systems) - *HYBRID*

Geico

🇺🇸United StatesJun 3, 2026, 4:26 AM UTC
Pluralis Research logo

Machine Learning Engineer - Distributed ML Systems

Pluralis Research

🌍Australia, United StatesMay 28, 2026, 4:07 AM UTC
Roblox logo

Principal Software Engineer - Creator Distributed Systems & Storage

Roblox

🇺🇸United StatesMay 28, 2026, 1:01 AM UTC
Institute of Foundation Models logo

Senior Distributed Systems Engineer

Institute of Foundation Models

🇺🇸United StatesMay 28, 2026, 12:07 AM UTC