Senior Storage Software Engineer – DGX Cloud
Posted 2hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Storage software engineer leading open-source distributed file systems for NVIDIA’s AI infrastructure. Troubleshooting massive GPU clusters and setting performance, durability, and tuning standards.
Responsibilities:
- Contribute code to open-source parallel and distributed file systems and distributed object storage
- Upstream fixes and features and engage with upstream communities and maintainers
- Write and review production code as a hands-on storage software lead
- Read kernel, NFS, NVMe-oF, or SPDK source to diagnose bugs
- Make final technical calls on storage deliveries against measurable targets
- Triage, troubleshoot, and root-cause complex storage issues across very large GPU clusters
- Investigate I/O and metadata performance, data corruption, and recovery
- Validate storage architecture, capabilities, performance, and durability
- Run scale tests, benchmarks, and recovery drills
- Qualify new builds against measurable performance and durability targets
- Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure
- Help operators and internal customers apply storage guidelines
- Collaborate with training, inference, accelerated-computing, SRE, operations, networking, and security teams
- Collaborate with cloud providers, neocloud operators, and storage vendors on common architecture
- Use modern AI coding and agentic tools to accelerate building, debugging, validation, and operations
Requirements:
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience
- Over 12 years of direct experience in storage software engineering
- Extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale
- Contributions to open-source projects involving a distributed or parallel file system
- Hands-on experience writing and reviewing production code, examining file system, kernel, NVMe-oF, or SPDK source, and conducting scale tests or recovery drills
- Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance
- Strong proficiency in at least one systems language: C, C++, Rust, or Go
- Proficiency in Python
- Comfortable in Linux kernel storage and networking stacks, including block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, and multipath
- Solid understanding of object storage, including S3 / Swift-class
- Solid understanding of block storage, including NVMe-oF and iSCSI
- Strong written and verbal communication
- Comfort operating in a 24/7 production environment
- Security-first approach
- Maintainers or sustained contributions to widely used public projects
- Experience crafting or operating storage for AI training or inference at very large GPU scale
- Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience
- Kubernetes and CSI driver development for storage
- Hands-on experience with SPDK, libfabric, or FUSE performance optimization
Benefits:
- Equity
- Benefits




















