NVIDIA's IT Storage Engineering team architects, designs, deploys, and manages petabyte-scale storage infrastructure that serves as the foundation for some of the most demanding workloads in the industry.

As the Senior Manager, Storage Engineering, you will lead a team of storage engineers and collaborate across NVIDIA's global IT organization to ensure that our storage infrastructure scales with the company's relentless pace of innovation.

Responsibilities:

  • Lead petabyte-scale storage deployments across NVIDIA's on-premises data centers and major cloud service providers (CSPs), owning the end-to-end lifecycle from design and procurement through physical installation, configuration, and production hand-off.
  • Engineer and maintain automation pipelines that integrate deployed storage systems with interdependent tooling, ensuring a single source of truth across the entire storage fleet.
  • Define and drive continuous improvement initiatives focused on data center efficiency, optimizing DC power consumption, and minimizing rack space footprint.
  • Develop and maintain self-service tools and dashboards that enable internal customers to track real-time storage capacity availability, consumption trends, and projected growth.
  • Coordinate and partner with partner infrastructure teams to present a unified capacity view, resolve cross-domain bottlenecks, and contribute to NVIDIA's holistic infrastructure capacity planning process.
  • Manage vendor relationships and lead hardware refresh and EOL planning cycles across the storage portfolio.
  • Recruit, mentor, and grow a team of storage deployment engineers; establish engineering standards, runbooks, and on-call practices to ensure a high operational bar.
  • Partner with architecture and security teams to evaluate new storage technologies, drive POCs, and translate findings into production-ready deployment standards.
  • Own and engineer the full hardware and data lifecycle, ensuring compliance, auditability, and zero unplanned data loss at every stage.
  • Define and govern data lifecycle management policies in close collaboration with key stakeholders across Chip Design and Software Engineering teams.

Requirements:

  • BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience.
  • 12+ years of overall experience in large-scale storage architecture, operations, production engineering, or infrastructure.
  • 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams.
  • Deep, protocol-level knowledge of enterprise storage systems spanning block, file, object, and high-performance parallel file systems.
  • Hands-on experience deploying and operating solutions from multiple major storage vendors.
  • Solid understanding of storage hardware internals, drive types and endurance profiles, controller architectures, shelf and enclosure design, cabling standards, and failure domain planning.
  • Working knowledge of bare-metal server hardware and data center network hardware relevant to storage connectivity.
  • In-depth expertise in observability tooling, building and maintaining Prometheus exporters, Grafana dashboards, and alerting rulesets.
  • Strong configuration management skills using Ansible for automated provisioning and day-2 operations of storage systems at scale.
  • Scripting and automation proficiency in Python and/or Bash; experience integrating with REST APIs.
  • Proven track record of managing large-scale, multi-vendor storage environments in a fast-paced, high-availability production setting.

Benefits:

  • Equity
  • Benefits