Neo Cloud - Principal AI Cloud Storage Engineer
- Hiring from
- United States
- Work type
- Hybrid
- Posted
Is this job info correct?
509,707 remote jobs, straight from company career pages
100% free · New jobs every hour
Show job descriptionHide job description
Member of Technical Staff, AI Cloud Storage
Location: Bay Area/ Seattle/ Remote
Reports to: CTO
About Neo Cloud
Neo Cloud is a fast growing next-generation AI neocloud, founded by ex NVIDIA, CoreWeave and Intel engineering leaders. We offer you the opportunity to build and operate next generation AI infrastructure that powers the world's most demanding AI workloads for the largest AI labs. We are rethinking how to run large scale AI infra for instance by leveraging agentic AI to build digital twins to de-risk our physical deployments. Our agents are deploying and monitoring our fleet to maximize the uptime of our infra.
About the Role
We're looking for a Principal Software Engineer to define and build Neo Cloud's AI cloud storage platform. This is a senior individual-contributor role for an experienced system architect who can operate at the intersection of customer needs, system architecture, and production operations — translating what AI/ML customers will need in the future, into a storage system that is performant, resilient, and economical at massive scale.
Key Responsibilities
Identifying customer requirements:
- Engage directly with customers, solutions architects, and product teams to understand future storage requirements for AI/ML workloads.
- Extrapolate from current usage patterns and industry trends to anticipate future requirements.
- Partner with product management to prioritize platform investments based on near-term customer needs and longer-term strategic bets.
System design, implementation, and operations:
- Design, implement, and operate AI cloud object and file storage systems, including the data path, metadata path, and control plane.
- Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions — including local node caching strategies, the network hardware and protocols that move data between storage and compute (e.g., RDMA, high-throughput NICs, congestion control), and the underlying storage hardware and software (media, erasure coding, metadata services) — and understanding how decisions in one layer constrain or unlock the others.
- Architect for the specific demands of AI workloads: very high aggregate throughput to keep GPU/accelerator clusters fed, support for massive numbers of small and large objects and files, efficient checkpointing at scale, and predictable tail latency under heavy concurrent load.
- Drive core storage system design decisions, including durability and consistency models, erasure coding and replication strategies, metadata scalability, multi-tenancy and isolation, and S3-compatible and POSIX/file-protocol API design.
- Take end-to-end ownership of services in production: build for observability and operability from day one, participate in on-call, lead incident response and root-cause analysis for critical issues, and drive long-term reliability and performance improvements.
- Identify and eliminate performance bottlenecks and scalability limits before they become customer-facing problems; lead capacity planning for rapid growth.
- Deep understanding of data privacy and security and its implication on performance.
- Ability to partner with network engineers to deliver complete AI storage system.
Technical leadership:
- Strong bias to action and resolution of technical decisions.
- Set technical direction and best practices for the storage organization; author and review design documents for significant architectural changes.
- Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
- Influence technical strategy across adjacent teams.
Qualifications
- 10+ years of professional software engineering experience building and operating cloud storage systems in production.
- Direct experience with AI-focused storage platforms such as DDN, Weka, or VAST, including their architectural approaches to throughput, caching, and GPU-cluster integration.
- Deep understanding of distributed systems fundamentals: consistency models, replication, consensus, partitioning, failure detection, and recovery.
- Proven experience operating high-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
- Strong systems programming skills (e.g., Go, C++, Rust, or Java) and comfort working across the stack from low-level I/O and networking to distributed control planes.
- Excellent written and verbal communication skills.
Nice to have
- Experience designing or tuning local node caching layers to accelerate AI training and inference data access.
- Experience with high-performance networking.
- Contributions to open-source storage projects, relevant patents, or published technical papers/talks.