NVIDIA's IT Storage Engineering team architects, designs, deploys, and manages petabyte-scale storage infrastructure that serves as the foundation for some of the most demanding workloads in the industry.
As the Senior Manager, Storage Engineering, you will lead a team of storage engineers and collaborate across NVIDIA's global IT organization to ensure that our storage infrastructure scales with the company's relentless pace of innovation.
Responsibilities:
- Lead petabyte-scale storage deployments across NVIDIA's on-premises data centers and major cloud service providers (CSPs), owning the end-to-end lifecycle from design and procurement through physical installation, configuration, and production hand-off.
- Engineer and maintain automation pipelines that integrate deployed storage systems with interdependent tooling, ensuring a single source of truth across the entire storage fleet.
- Define and drive continuous improvement initiatives focused on data center efficiency, optimizing DC power consumption, and minimizing rack space footprint.
- Develop and maintain self-service tools and dashboards that enable internal customers to track real-time storage capacity availability, consumption trends, and projected growth.
- Coordinate and partner with partner infrastructure teams to present a unified capacity view, resolve cross-domain bottlenecks, and contribute to NVIDIA's holistic infrastructure capacity planning process.
- Manage vendor relationships and lead hardware refresh and EOL planning cycles across the storage portfolio.
- Recruit, mentor, and grow a team of storage deployment engineers; establish engineering standards, runbooks, and on-call practices to ensure a high operational bar.
- Partner with architecture and security teams to evaluate new storage technologies, drive POCs, and translate findings into production-ready deployment standards.
- Own and engineer the full hardware and data lifecycle, ensuring compliance, auditability, and zero unplanned data loss at every stage.
- Define and govern data lifecycle management policies in close collaboration with key stakeholders across Chip Design and Software Engineering teams.
Requirements:
- BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
- 12+ years of overall experience in large-scale storage architecture, operations, production engineering, or infrastructure.
- 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
- Deep, protocol-level knowledge of enterprise storage systems spanning block, file, object, and high-performance parallel file systems.
- Hands-on experience deploying and operating solutions from multiple major storage vendors.
- Solid understanding of storage hardware internals, drive types and endurance profiles, controller architectures, shelf and enclosure design, cabling standards, and failure domain planning.
- Working knowledge of bare-metal server hardware and data center network hardware relevant to storage connectivity.
- In-depth expertise in observability tooling, building and maintaining Prometheus exporters, Grafana dashboards, and alerting rulesets.
- Strong configuration management skills using Ansible for automated provisioning and day-2 operations of storage systems at scale.
- Scripting and automation proficiency in Python and/or Bash; experience integrating with REST APIs.
- Proven track record of managing large-scale, multi-vendor storage environments in a fast-paced, high-availability production setting.
Benefits:
- Equity
- Benefits

