Senior Storage Software Engineer – DGX Cloud

Posted 2hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Storage software engineer leading open-source distributed file systems for NVIDIA’s AI infrastructure. Troubleshooting massive GPU clusters and setting performance, durability, and tuning standards.

Responsibilities:

  • Contribute code to open-source parallel and distributed file systems and distributed object storage
  • Upstream fixes and features and engage with upstream communities and maintainers
  • Write and review production code as a hands-on storage software lead
  • Read kernel, NFS, NVMe-oF, or SPDK source to diagnose bugs
  • Make final technical calls on storage deliveries against measurable targets
  • Triage, troubleshoot, and root-cause complex storage issues across very large GPU clusters
  • Investigate I/O and metadata performance, data corruption, and recovery
  • Validate storage architecture, capabilities, performance, and durability
  • Run scale tests, benchmarks, and recovery drills
  • Qualify new builds against measurable performance and durability targets
  • Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure
  • Help operators and internal customers apply storage guidelines
  • Collaborate with training, inference, accelerated-computing, SRE, operations, networking, and security teams
  • Collaborate with cloud providers, neocloud operators, and storage vendors on common architecture
  • Use modern AI coding and agentic tools to accelerate building, debugging, validation, and operations

Requirements:

  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience
  • Over 12 years of direct experience in storage software engineering
  • Extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale
  • Contributions to open-source projects involving a distributed or parallel file system
  • Hands-on experience writing and reviewing production code, examining file system, kernel, NVMe-oF, or SPDK source, and conducting scale tests or recovery drills
  • Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance
  • Strong proficiency in at least one systems language: C, C++, Rust, or Go
  • Proficiency in Python
  • Comfortable in Linux kernel storage and networking stacks, including block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, and multipath
  • Solid understanding of object storage, including S3 / Swift-class
  • Solid understanding of block storage, including NVMe-oF and iSCSI
  • Strong written and verbal communication
  • Comfort operating in a 24/7 production environment
  • Security-first approach
  • Maintainers or sustained contributions to widely used public projects
  • Experience crafting or operating storage for AI training or inference at very large GPU scale
  • Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience
  • Kubernetes and CSI driver development for storage
  • Hands-on experience with SPDK, libfabric, or FUSE performance optimization

Benefits:

  • Equity
  • Benefits