Senior Storage Software Engineer - DGX Cloud

Other Jobs To Apply

No other job posts for this day.

NVIDIA DGXC Storage team handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential, and between launching a frontier model on time or missing the deadline by months. We’re looking for a hands-on Storage Software Engineer to join the storage team as an individual contributor and technical lead. You will contribute to open-source parallel and distributed file systems and keep our largest GPU clusters fast, reliable, and durable. You will stay deeply hands-on: writing and reviewing production code, chasing root causes in the field, and setting the configuration and tuning standards our GPU fleets run on. This is a chance to do foundational storage engineering for the AI era at the company that introduced accelerated computing. What you’ll be doing: * Contribute to open-source file systems. Contribute code to open-source parallel and distributed file systems, and distributed object storage. Upstream fixes and features, and engage directly with the upstream communities and maintainers. * Serve as a hands-on storage software lead. Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it. Make the final technical calls on storage deliveries against measurable targets. * Triage and troubleshoot at scale. Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery. * Validate architecture and capabilities. Validate storage architecture, capabilities, performance, and durability. Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets. * Recommend configuration, tuning, and guidelines. Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them. * Partner broadly. Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture. * Work AI-first. Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations. What we need to see: * BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience. Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale. * Contributions to open-source projects involving a distributed or parallel file system. You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale tests or recovery drills instead of assigning them to others. * Experience diagnosing and resolving storage problems in extensive GPU or HPC clusters, including analysis of I/O and metadata performance. * Strong proficiency in at least one systems language (C, C++, Rust, or Go) and proficiency in Python; comfortable in the Linux kernel storage and networking stacks (block layer, RDMA / RoCE / InfiniBand, NVMe, page cache, VFS, multipath). * Solid understanding of object storage (S3 / Swift-class) and block storage (NVMe-oF, iSCSI). * Strong written and verbal communication; capable of clarifying complex technical trade-offs to engineers, SREs, vendors, and internal customers. * Comfort operating in a 24/7 production environment where storage incidents directly impact GPU availability, with a security-first approach baked into every build. * 100% hands-on engineering. You write and review production code, read file system, kernel, NVMe-oF, or SPDK source to chase bugs, and run scale tests or recovery drills yourself rather than delegating. Ways to stand out from the crowd: * Maintainers or sustained contributions to widely used public projects. * Experience crafting or operating storage for AI training or inference at very large GPU scale, with measurable gains in GPU utilization or reductions in I/O bottlenecks. * Kernel and file system development experience, metadata scalability, data placement, failure recovery, or HSM or equivalent experience. * Kubernetes and CSI driver development for storage. * Hands-on experience with SPDK, libfabric, or FUSE performance optimization. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. Our invention serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is seeking exceptional individuals like you to help us drive the next wave of artificial intelligence. NVIDIA is widely considered one of the world's most desirable employers in technology. We have some of the world's most forward-thinking and passionate people working for us. If you're creative and autonomous, we want to hear from you! Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6. You will also be eligible for equity and benefits. Applications for this job will be accepted at least until September 24, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Back to blog

Common Interview Questions And Answers

1. HOW DO YOU PLAN YOUR DAY?

This is what this question poses: When do you focus and start working seriously? What are the hours you work optimally? Are you a night owl? A morning bird? Remote teams can be made up of people working on different shifts and around the world, so you won't necessarily be stuck in the 9-5 schedule if it's not for you...

2. HOW DO YOU USE THE DIFFERENT COMMUNICATION TOOLS IN DIFFERENT SITUATIONS?

When you're working on a remote team, there's no way to chat in the hallway between meetings or catch up on the latest project during an office carpool. Therefore, virtual communication will be absolutely essential to get your work done...

3. WHAT IS "WORKING REMOTE" REALLY FOR YOU?

Many people want to work remotely because of the flexibility it allows. You can work anywhere and at any time of the day...

4. WHAT DO YOU NEED IN YOUR PHYSICAL WORKSPACE TO SUCCEED IN YOUR WORK?

With this question, companies are looking to see what equipment they may need to provide you with and to verify how aware you are of what remote working could mean for you physically and logistically...

5. HOW DO YOU PROCESS INFORMATION?

Several years ago, I was working in a team to plan a big event. My supervisor made us all work as a team before the big day. One of our activities has been to find out how each of us processes information...

6. HOW DO YOU MANAGE THE CALENDAR AND THE PROGRAM? WHICH APPLICATIONS / SYSTEM DO YOU USE?

Or you may receive even more specific questions, such as: What's on your calendar? Do you plan blocks of time to do certain types of work? Do you have an open calendar that everyone can see?...

7. HOW DO YOU ORGANIZE FILES, LINKS, AND TABS ON YOUR COMPUTER?

Just like your schedule, how you track files and other information is very important. After all, everything is digital!...

8. HOW TO PRIORITIZE WORK?

The day I watched Marie Forleo's film separating the important from the urgent, my life changed. Not all remote jobs start fast, but most of them are...

9. HOW DO YOU PREPARE FOR A MEETING AND PREPARE A MEETING? WHAT DO YOU SEE HAPPENING DURING THE MEETING?

Just as communication is essential when working remotely, so is organization. Because you won't have those opportunities in the elevator or a casual conversation in the lunchroom, you should take advantage of the little time you have in a video or phone conference...

10. HOW DO YOU USE TECHNOLOGY ON A DAILY BASIS, IN YOUR WORK AND FOR YOUR PLEASURE?

This is a great question because it shows your comfort level with technology, which is very important for a remote worker because you will be working with technology over time...

Other Jobs To Apply

Remote Pharmacy Tech: Patient Intake & Data Entry

Remote 3rd Shift Customer Service Representative - Financial Services | Overnight Support Specialist (11PM-7AM) with Training

Cart Associate - JFK John F. Kennedy Airport- Part Time

Overnight Crew Member10PM-6AM Starting at $14 an Hour Same Day Pay Options

Customer Service/Sales – Bilingual Spanish

Utilization Management Nurse Consultant

Part-Time Airport Customer Assistant & Kiosk Specialist

Social Media Content Specialist - Remote Work

Utilization Management (UM) Nurse Consultant – Work at Home in EST or CST Zone

Executive Assistant (100% Remote)

Customer Service Associate I

Amazon Delivery Associate

Paid Summer Sales Internship - No Experience + Training

Warehouse Associate/Shipping Specialist - Kearny Mesa

Casting Assistant

Online Product Tester - Get Paid for Reviews

Amazon Delivery Driver - $26/Hour + Benefits + Bonuses

Imaging Analyst - Remote based in the US

Remote Overnight Schedules – Detail-Oriented Night reputed company $25–$35/Hour (No Degree Needed)

Retail Stocking Team Lead - Part-Time

Security Officer

Seasonal Full Time Hourly Warehouse Operations Openings (T3865)

Amazon Package Delivery Driver - Earn $15.00 - $28.50/hr

Store Team Member

Distribution Center Supervisor

amazon flex $15+/ Hour (Sign on Bonus)!

Amazon Remote Jobs (No Degree, Part/Full Time, Customer Support) ? Entry Level

Customer Service Representative (remote)

Warehouse Material Handler Part Time 3rd Shift

Remote Data Entry Clerk - Typing - Part Time Entry Level-United States

Passenger Service Agent (BWI)

Call Center Agent (PT or FT - 100% Work From Home)

Store Associate - Davis (00127)

Full Time Housekeeping

WORK FROM HOME POSITION - START THIS WEEK | FLEXIBLE HOURS | IMMEDIATE OPENINGS

Sales Representative - Remote - Ontario, California

Account Manager - Texas - Remote Work

Cook - Mercer Atlanta

Utilization Management Nurse

Remote RN Telephone Triage - $1000 Sign On Bonus!

Driver Merchandiser Assistant at Great Lakes Coca-Cola Alsip, IL

Security Officer Full Time Assisted Living Facility

Barista

Remote Work From Home Data Entry Jobs $1400 Weekly

Amazon Flex Delivery Driver

Warehouse Associate - Pallet Repair

Warehouse Loading Associate - Now Hiring

QA Requirements Tester - Activision

Restaurant Team Member

Permit Operations Lead (Fully Remote)