Cotidal logo
Moves you to
United States
Support
Visa sponsorshipRelocation support
Posted
Sep 25, 2026
Is this job info correct?

Who We Are

Cotidal is building endless compute for humanity.

We believe the current compute scarcity is a structural and enduring trend, not a passing shortage: demand for AI will outpace the world's ability to scale infrastructure for decades to come, and no single supplier or architecture will meet it. Cotidal is an AI infrastructure company that designs, builds, and operates the accelerator fleets ambitious AI teams train and serve on — expanding the supply of usable compute, and working toward a world where the cost of compute stops deciding which ideas get tried.

We believe the future of AI compute is heterogeneous, and we are building the platform for that future — designed to take on new silicon as it matures. Our first clusters are committed and come online this year, and we work directly with our initial design partners, so every layer, from the data center to the customer API, is being built by the people you will sit next to.

We’ve raised two rounds of funding in our first three months, led by top AI and semiconductor funds, with strategic financial support from partners across the chip supply chain.

The Role

You’ll own the hardware systems and cluster architecture behind Cotidal’s compute platform. You’ll turn workload requirements into hardware specifications, capacity plans, and deployable cluster designs, making decisions across performance, reliability, power, cost, and availability.

Working directly with customers, teammates, and hardware partners, you’ll define how our systems fit together and how they scale. You’ll lead hardware selection and architectural decisions from the first design through validation and deployment, and use production evidence to improve the next generation.

What You Will Do

  • Hardware Systems & Cluster Architecture: Design and evaluate the systems that connect accelerators, CPUs, memory, networking, and storage into high-performance clusters. Review vendor reference designs, validate compatibility and sizing assumptions, and define the configurations and interfaces needed for deployment. Own server configurations, rack and pod layouts, and failure domains, working with network and data-center engineers on connectivity, power, cooling, and serviceability.

  • Capacity Planning & Expansion: Translate customer demand and workload profiles into hardware capacity requirements and expansion plans. Size compute, memory, storage, and network capacity; quantify usable capacity and growth headroom; and identify where bottlenecks limit the system. Work with platform, supply chain, and data-center teams to align configurations and deployment phases with delivery lead times and site readiness.

  • Hardware Platforms & Vendor Engineering: Evaluate new hardware, compare configurations, and own technical specifications and engineering approval of bills of materials and substitutions. Work directly with vendors to resolve compatibility issues, close gaps against system requirements, and bring new platforms into deployment.

  • Design Validation & Improvement: Own hardware design acceptance criteria and work with platform engineers to validate configurations under representative workloads and failure conditions. Investigate problems across hardware, firmware, and system interfaces. Drive corrective changes with vendors and subsystem owners, and carry lessons from deployment and operation into future designs.

What We Are Looking For

  • You’ve led hardware architecture and system design for production compute infrastructure, owned the critical engineering tradeoffs, and carried designs through validation and deployment. You’ve remained accountable for design issues discovered in production.

  • You understand how accelerators, CPUs, memory, interconnects, networking, and storage affect a complete system. You can turn workload requirements into quantitative sizing, capacity plans, and engineering tradeoffs.

  • You’ve evaluated hardware configurations and worked directly with vendors on technical requirements, compatibility, and issue resolution.

  • You stay close to the hardware. You can use Linux, scripting, benchmarks, and failure data to investigate problems and distinguish hardware limitations from configuration or software issues.

  • You turn open-ended requirements into a buildable design, make tradeoffs clear, and follow through when validation or production evidence challenges your assumptions.

Especially Valuable

Experience in one or more of the following:

  • Accelerator systems, server and rack architecture, and scaling compute across multiple racks.

  • OEM/ODM collaboration, component qualification, and introducing new hardware platforms.

  • High-performance interconnects and storage architectures for distributed workloads, including RDMA networks and parallel file systems.

  • Rack integration, power and thermal constraints, liquid-cooling interfaces, and serviceability.

Who You Will Work With

Cotidal’s founding team brings together repeat founders, engineers, and operators with backgrounds at xAI, Tesla, Google, Microsoft, SSI, Figure AI, and Cursor. Your teammates have built AI infrastructure that much of the industry serves models on—and you’ll work directly alongside them.

In our first three months, we secured chip allocations, locked in our first site, and signed our first design partners. You’ll join while the architecture, engineering culture, and team are still taking shape. The systems you build, the standards you set, and the people you help recruit will shape what Cotidal becomes.

Location & What We Offer

Location: Palo Alto, in person.

Compensation: Competitive salary and equity.

Visa sponsorship: We sponsor work visas and help you navigate the process.

Benefits: Medical, dental, and vision coverage, unlimited PTO, and relocation support as needed.

Similar jobs

Apply for this job