CO

Software Engineer, Compute Platform

Moves you to
United States
Support
Visa sponsorshipRelocation support
Posted
Sep 24, 2026
Is this job info correct?

Who We Are

Cotidal is building endless compute for humanity.

We believe the current compute scarcity is a structural and enduring trend, not a passing shortage: demand for AI will outpace the world's ability to scale infrastructure for decades to come, and no single supplier or architecture will meet it. Cotidal is an AI infrastructure company that designs, builds, and operates the accelerator fleets ambitious AI teams train and serve on — expanding the supply of usable compute, and working toward a world where the cost of compute stops deciding which ideas get tried.

We believe the future of AI compute is heterogeneous, and we are building the platform for that future — designed to take on new silicon as it matures. Our first clusters are committed and come online this year, and we work directly with our initial design partners, so every layer, from the data center to the customer API, is being built by the people you will sit next to.

We’ve raised two rounds of funding in our first three months, led by top AI and semiconductor funds, with strategic financial support from partners across the chip supply chain.

The Role

You’ll define and build Cotidal’s compute platform, from the software that brings accelerator capacity online to the services AI teams use to provision resources and run workloads. Customers build on this platform directly: its APIs, reliability, and cost are what they experience as Cotidal. You’ll make architectural decisions across provisioning, scheduling, and fleet reliability, carrying systems from their first implementation through production scale.

Working directly with customers and teammates, you’ll identify the problems that matter, set technical priorities, and decide how to solve them. We’re looking for engineers with depth in one area and the judgment and curiosity to work across services, operating systems, networking, storage, and hardware.

What You Will Do

Your initial focus will reflect your experience and the team’s needs, with room to take on broader ownership as the platform grows.

Compute Platform & Control Plane: Design and build the APIs, resource models, and distributed services that manage capacity and orchestrate workloads. Own allocation, scheduling, and resource lifecycle, including tenant isolation and safe resource reassignment. Make deliberate tradeoffs between utilization, reliability, and the customer experience as demand and fleet size grow.

Cluster Bring-up & Qualification: Own bring-up of new capacity, from installed hardware to clusters customers run on. Build the provisioning, configuration, and qualification software across compute, networking, and storage, with repeatable tests for health, performance, and recovery, and define and verify the criteria for releasing each cluster into production. Partner with the HPC engineers who design our clusters: they bring hardware depth to bring-up, and you bring workload requirements back into the next design.

Fleet Lifecycle & Reliability: Define how the fleet is observed, maintained, and recovered. Build health checks, safe upgrade and rollback paths, failure isolation, and automated remediation. Own production diagnosis and incident response for the systems you build, turning failures into software improvements that reduce recovery time, make the fleet cheaper to run, and keep more capacity available for useful work.

What We Are Looking For

  • You’ve designed, built, and operated production infrastructure software—services, controllers, or systems that manage real resources. You’ve carried it through failures, growth, and changing requirements.

  • You have deep expertise in a part of the stack—distributed control planes, operating systems, networking, storage, or cluster infrastructure—and follow difficult problems across its boundaries. You work from source code, system behavior, and measurements to establish a cause and verify a fix.

  • You can reason about distributed state, concurrency, and partial failure. You design systems that recover interrupted operations, reconcile changes safely, and make their behavior understandable to the people operating them.

  • You identify consequential problems before someone writes a specification for them. You work directly with users, decide what is worth building, and carry it through to a reliable system. You make your reasoning clear and change your approach when the evidence calls for it.

Especially Valuable

Experience in one or more of the following:

  • Building infrastructure control planes, including resource APIs, state reconciliation, and workflows that recover interrupted operations.

  • Developing scheduling and resource allocation systems that account for topology, scarce capacity, quotas, and tenant isolation, using Kubernetes, Slurm, or similar systems.

  • Building provisioning and configuration systems that bring new capacity online, including host images, firmware, drivers, network boot, and hardware management through tools and protocols such as PXE and Redfish.

  • Building fleet health and recovery systems, including capacity qualification, staged upgrades, failure isolation, and automated repair that verifies recovery before returning resources to service.

  • Diagnosing reliability and performance problems across Linux, networking, storage, and distributed workloads, including accelerator clusters, RDMA networks, or distributed file systems.

Who You Will Work With

Cotidal’s founding team brings together repeat founders, engineers, and operators with backgrounds at xAI, Tesla, Google, Microsoft, SSI, Figure AI, and Cursor. Your teammates have built AI infrastructure that much of the industry serves models on—and you’ll work directly alongside them.

In our first three months, we secured chip allocations, locked in our first site, and signed our first design partners. You’ll join while the architecture, engineering culture, and team are still taking shape. The systems you build, the standards you set, and the people you help recruit will shape what Cotidal becomes.

Location & What We Offer

Location: Palo Alto, in person.

Compensation: Competitive salary and equity.

Visa sponsorship: We sponsor work visas and help you navigate the process.

Benefits: Medical, dental, and vision coverage, unlimited PTO, and relocation support as needed.

Similar jobs

Apply for this job