CO

Software Engineer, AI Systems & Performance

Moves you to
United States
Support
Visa sponsorshipRelocation support
Posted
Sep 24, 2026
Is this job info correct?

Who We Are

Cotidal is building endless compute for humanity.

We believe the current compute scarcity is a structural and enduring trend, not a passing shortage: demand for AI will outpace the world's ability to scale infrastructure for decades to come, and no single supplier or architecture will meet it. Cotidal is an AI infrastructure company that designs, builds, and operates the accelerator fleets ambitious AI teams train and serve on — expanding the supply of usable compute, and working toward a world where the cost of compute stops deciding which ideas get tried.

We believe the future of AI compute is heterogeneous, and we are building the platform for that future — designed to take on new silicon as it matures. Our first clusters are committed and come online this year, and we work directly with our initial design partners, so every layer, from the data center to the customer API, is being built by the people you will sit next to.

We’ve raised two rounds of funding in our first three months, led by top AI and semiconductor funds, with strategic financial support from partners across the chip supply chain.

The Role

You’ll define and build Cotidal’s workload software stack for training, reinforcement learning, and inference. You’ll make architectural decisions and carry systems from their first implementation through production scale as models, workloads, and accelerator architectures evolve.

Working directly with customers and teammates, you’ll identify the problems that matter, set technical priorities, and decide how to solve them. We’re looking for engineers with depth in one area, the judgment to navigate open-ended problems, and the curiosity to work across the stack.

What You Will Do

Model & Framework Integration: Design and build the framework integrations and distributed execution systems that bring training, RL, and inference workloads to Cotidal’s accelerators. Extend ML frameworks and libraries, resolve compatibility and compilation issues, and shape support for new model architectures and workflows.

Performance Engineering: Set performance goals, build representative benchmarks, and identify the changes that will have the greatest impact on real workloads. Improve model code, compilers, runtimes, and kernels to increase training efficiency, shorten RL iteration time, and improve inference throughput and latency while preserving correctness and model quality.

Production Reliability: Own the reliability of the systems you build as workloads scale. Define validation and regression tests, observability, checkpointing, and recovery mechanisms. Turn production experience into reusable capabilities and upstream improvements that make new workloads easier to support and existing ones easier to operate.

What We Are Looking For

  • You’ve built software for training, inference, or accelerated computing and stayed close to how it behaves in practice—through performance limits, failures, and changing requirements.

  • You follow a problem wherever it leads: from model code and distributed execution into the runtime or hardware, then back to whether the improvement helps the overall workload.

  • You have strong software engineering fundamentals and understand how compute, memory, and communication shape system performance. You’re comfortable reading unfamiliar code and using experiments to test your understanding.

  • You turn open-ended goals into a technical plan, decide what to prioritize, and carry the work through to a reliable system. You take responsibility for the outcome and adjust your approach as you learn.

  • You work directly with users and teammates, make your reasoning clear, and seek out the expertise needed to move a problem forward. You have opinions about how systems should be built and have changed your mind when the evidence called for it.

Especially Valuable

Experience in one or more of the following:

  • ML frameworks and compilers such as PyTorch, JAX, or XLA, and hands-on work with GPUs, TPUs, Trainium, or other accelerator architectures.

  • RL and post-training systems, including rollout generation, weight synchronization, and frameworks such as verl.

  • Distributed training and execution, including sharding, model parallelism, collective communication, and overlapping computation with communication.

  • Kernel development or compiler optimization using Pallas, Triton, CUDA, or similar tools.

  • Inference optimization with vLLM, SGLang, or similar systems, including batching, KV caches, and quantization.

Who You Will Work With

Cotidal’s founding team brings together repeat founders, engineers, and operators with backgrounds at xAI, Tesla, Google, Microsoft, SSI, Figure AI, and Cursor. Your teammates have built AI infrastructure that much of the industry serves models on—and you’ll work directly alongside them.

In our first three months, we secured chip allocations, locked in our first site, and signed our first design partners. You’ll join while the architecture, engineering culture, and team are still taking shape. The systems you build, the standards you set, and the people you help recruit will shape what Cotidal becomes.

Location & What We Offer

Location: Palo Alto, in person.

Compensation: Competitive salary and equity.

Visa sponsorship: We sponsor work visas and help you navigate the process.

Benefits: Medical, dental, and vision coverage, unlimited PTO, and relocation support as needed.

Similar jobs

Apply for this job