Node Systems Lead
- Moves you to
- United States
- Support
- Relocation support
- Posted
Show job descriptionHide job description
About Etched
Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.
Job Summary
We’re hiring a Node Systems Lead to join our Supercomputing organization. This team builds the software that enables Etched to deploy inference clusters at gigawatt-scale. This role presents an opportunity to shape how frontier inference hardware is configured and managed for our customers.
We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads. Node Systems owns the software layer of each individual Etched node, from host software and system configuration to the interfaces between rack components. It tunes CPU scheduling, memory, networking, and host-to-accelerator communication so models get every bit of performance and reliability our hardware can deliver, then turns those gains into tested software and configurations that ship in every system.
We’re looking for a leader who can set the technical direction and dive deep into difficult systems problems. Someone who has taken hardware from early bring up into production, can build the best team in the industry, and wants to stay close to architecture and code. The team’s work is central in enabling Etched to deploy inference clusters at massive scale.
Key Responsibilities
Lead and develop the Node Systems team, setting priorities and giving engineers clear ownership.
Set the technical direction for host software, system configuration and rack component interfaces.
Guide the team’s performance work across CPU scheduling, memory, networking and host-to-accelerator communication.
Ensure system optimizations become tested, maintainable software and configurations ready for deployment.
Lead the development of rack simulation and diagnostic capabilities that support platform development and debugging.
Partner with inference software, firmware, and hardware teams to architect our system design for current and next-gen products
Align with Fleet Software on the configurations and interfaces needed to manage and monitor deployed systems at the cluster-level
Work with manufacturing and test engineering to establish the software baseline and test coverage needed to ship reliable systems.
Stay close to the implementation through design reviews, code contributions and hands-on debugging.
You may be a good fit if you have (Must-have qualifications)
Strong experience developing and debugging production systems software in C, C++ or Rust on Linux.
Experience leading technical projects and mentoring engineers while staying hands on.
Deep understanding of operating systems fundamentals, including scheduling, concurrency, memory management and I/O.
Demonstrated ability to investigate performance problems on real hardware and validate improvements through profiling and measurement.
Experience taking system software or performance improvements from prototype through production deployment.
Understanding of hardware/software interactions and the ability to debug across application, kernel and device boundaries.
Ability to work closely with hardware and software teams and translate workload requirements into concrete system changes.
Strong candidates may also have experience with (Nice-to-have qualifications)
Experience with low-latency systems, high-frequency trading, HPC, or accelerator-based compute platforms.
Experience with Linux kernel development or debugging, CPU isolation, NUMA, interrupt affinity, and performance tuning.
Familiarity with PCIe, DMA, RDMA, device drivers, or high-performance networking.
Experience bringing up new hardware platforms or working closely with firmware and hardware engineers.
Exposure to manufacturing diagnostics, factory testing, or production system validation.
Experience building systems or engineering teams at an early-stage startup.
Benefits
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2,500 per month for those living within walking distance of the office
Relocation support for those moving to San Jose (Santana Row)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification
How we’re different
Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system, betting early on transformer and transformer-like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.
We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.