StackYak logo

Senior Compute Infrastructure Engineer

Hiring from
Worldwide
Work type
Remote
Posted
Oct 1, 2026
Is this job info correct?
  1. About StackYak

StackYak is building the infrastructure layer for AI.

We are an early-stage, funded company building software that brings compute, GPU infrastructure, networking, and inference together into one product. The opportunity is large, the market is moving quickly, and we are building for real production workloads from the start.

This is not internal IT. This is not a slow-moving infrastructure team maintaining someone else's platform. The infrastructure is the product.

We are a small, senior team with very little bureaucracy. This is a founding role in its discipline. You will be the first person here whose primary responsibility is the systems layer beneath the product, and the shape it takes will largely be yours to decide.

Treat this document as a starting point rather than a boundary. The people who do well here take ground early and are not asked to give it back.

We move quickly. We do not have months for someone to learn the fundamentals of their discipline. You should already be very good at what you do, be able to ramp into adjacent areas quickly, and be comfortable operating without perfect requirements or neatly defined boundaries.

The Role

We need someone who knows what stands between GPU capacity and production infrastructure, and can build it.

That capacity does not arrive in one form. Some of it comes as hardware, with everything that implies: firmware, drivers, hosts, and physical work carried out by people you will never meet in buildings you will rarely visit. Some of it comes from neoclouds and other providers, where you control the software and very little else, on terms set by someone whose interests are not yours.

The interesting problem is neither of those on its own. It is the gap between them — how much of it can honestly be abstracted away, how much has to stay visible, and how much the rest of the company should ever have to think about. Those are open questions here, and the person in this role is the one who gets to answer them.

You should be comfortable at both ends. On one side: Linux, kernel, firmware, drivers, virtualization, host hardening, and debugging a machine you cannot walk up to. On the other: provider APIs, capacity that is not permanent, and automation at a scale where nothing gets configured by hand.

Workloads will run on bare metal, in VMs, and in containers. The right answer will not be the same one twice, and choosing is part of the job.

This is not an architecture-only role. You will design systems and then build, debug, and operate them.

What You Will Own

Not tasks. Outcomes, and the authority that comes with them.

  • What the fleet is. What we own versus what we rent, where we standardise and where we deliberately do not, and what it costs us either way.

  • The machines themselves. Linux, kernel, firmware, drivers, CUDA and/or ROCm — down to the layer where the answer lives in a changelog rather than documentation.

  • Bare metal, VM, or container. You own the boundary between them, and the uncomfortable fact that the right answer will not be the same one twice.

  • Isolation. What we claim about keeping customers and workloads apart, and whether that claim survives someone attacking it.

  • Provisioning and recovery. A machine goes from nothing to serving, and from broken back to serving, without you personally being involved.

  • Provider integration. Neoclouds and other compute vendors: what each of them actually gives you as opposed to what they advertise, and the seams that difference creates.

  • Physical reality. Specifying hardware, and getting it from a purchase order to a booting machine through vendors, data-center partners, and remote hands you have never met.

  • Failures that cross layers, where the cause could be hardware, firmware, OS, virtualization, storage, or network, and nobody else wants to own the question.

  • Making it operable by other people, so the fleet does not quietly become a collection of one-off machines only you understand.

What Success Looks Like

You can be handed GPU capacity — some of it ours, some of it rented, none of it uniform — along with a set of business and security requirements, and turn it into production infrastructure the rest of the company can trust and operate.

You will help us answer questions such as:

  • Should this workload run on bare metal, a VM, or a container?

  • How should we provision, rebuild, and recover machines we cannot walk up to?

  • Where does owned hardware genuinely beat rented capacity, and where are we just paying for the privilege?

  • What should customer isolation look like?

  • Where should we standardize, and where should we preserve flexibility?

  • How do we operate owned hardware and external compute as one coherent platform rather than two?

  • Which parts of the infrastructure should become product capabilities rather than internal operations?

What We Need

  • The bar is what you have already done, not what you could learn. You should have done most of this:

  • Run production Linux at a depth where the answer was in the kernel, not in the docs.

  • Operated GPU infrastructure for AI, HPC, cloud, or similarly demanding workloads.

  • Carried NVIDIA and/or AMD platforms through a driver or firmware upgrade that did not go smoothly.

  • Provisioned bare metal at a count where configuring machines by hand stopped being an option.

  • Managed machines you could not physically reach, and recovered one anyway.

  • Chosen between bare metal, KVM, and containers with a real consequence attached, and been able to say why.

  • Written Terraform and Python that other people depend on in production — tooling, not glue.

  • Hardened hosts and isolated workloads for customers who were not permitted to see each other.

  • Debugged a failure that could have been hardware, firmware, OS, driver, network, or workload, and established which it was.

  • Made an infrastructure decision on bad information and then lived with it long enough to learn whether you were right.

You Will Be Especially Strong If

  • You have built or operated infrastructure at a neocloud, hyperscaler, GPU cloud, HPC environment, hosting company, data-center operator, or AI infrastructure company.

  • You have run a fleet that mixed owned hardware with rented capacity, and have opinions about what that costs you.

  • You have worked with multi-GPU and multi-node systems.

  • You understand NUMA, PCIe topology, RDMA, NIC placement, and why physical topology matters for AI workloads.

  • You have specified or accepted hardware, and know what goes wrong between the purchase order and a machine that boots.

  • You have designed infrastructure meant to serve multiple customers safely.

  • You have done this outside the confines of a tightly siloed enterprise team.

  • You build things on the side because systems are interesting to you, not only because an employer assigned a ticket.

This Is Probably Not For You If

  • Your idea of infrastructure ends at Kubernetes manifests or CI/CD pipelines.

  • You mostly operate through tickets and escalations.

  • You prefer designing systems that someone else implements.

  • You require clearly bounded ownership before touching a problem.

  • You have cloud experience but little understanding of what happens underneath the VM.

  • You know GPU APIs but have never been responsible for the machines underneath them.

  • You want to work only on hardware you own and consider provider-managed capacity beneath you — or the reverse, and would rather never think about a physical machine again.

  • You want months to become productive in the core areas of the role.

How We Work

  • Small, senior team with direct access to the founders.

  • Strong opinions, loosely held.

  • Everyone is expected to participate in technical decisions.

  • Everyone shares responsibility for production and on-call.

  • We value people who can move between design, implementation, debugging, and operations.

  • We care much more about what you have built and operated than degrees, certifications, or academic credentials.

  • We expect people to leave ego at the door, argue the technical case, make a decision, and then execute.

  • We are remote and distributed across time zones. Whether a role is an employment or a contract engagement depends on where you are, and we work that out at offer.

  • Hiring here is a few real conversations with the people you would actually work with, not a recruiter screen followed by a panel of strangers.

  • We are hiring across inference, infrastructure, and networking. The boundaries between the three are blurry on purpose. If you sit between two of them, say so.

  • This is an early-stage startup. The pace is high, the problems are hard, and the scope will change as we grow.

Compensation

Competitive compensation plus meaningful equity. Exact structure will depend on location, engagement model, and experience.

A Note For Agencies

We are not using external recruiters or agencies for this role, and we will not be persuaded otherwise by an email. We do not want your spam. We will not read the CVs you send, we will not reply to your follow-up, and no candidate you put in front of us creates a fee obligation of any kind. Do not contact us.

Similar jobs

Apply for this job