NVIDIA logo
NVIDIAPosted 4 weeks ago

Senior Manager, Storage Production Engineering

$272,000–$431,250 year

RemoteUnited States

Full TimeSenior LevelMasters DegreeEnterprise

Job Summary

Lead a team of Storage Production Engineers, designing and deploying large-scale storage systems including distributed storage, parallel file systems, and object storage. Drive reliability and efficiency by implementing automation, monitoring, and analytics for storage services. Own capacity planning, data lifecycle management, and high availability strategies while guiding incident response and root cause analysis. Partner with engineering and AI teams to optimize data pipelines and workflow performance. Manage on-call rotations, troubleshooting, and operational KPIs for production systems.

Required Qualifications

  • BS or MS in Computer Science, Storage Systems, or a related technical field, or equivalent experience
  • 12+ overall years of experience in large scale storage architecture, operations, production engineering, or infrastructure
  • 6+ years of people management or technical leadership experience with storage, infrastructure, or site reliability teams
  • Direct experience managing infrastructure operations including on-call rotations, incident response, ongoing maintenance, troubleshooting, and optimization of production systems, managing SLOs and operational KPIs
  • Hands on experience with parallel file systems (such as Lustre or GPFS), distributed storage (such as Ceph or MinIO), and enterprise object or NAS platforms (such as S3 compatible systems, NetApp, or Pure Storage)
  • Strong knowledge of block, file, and object storage, including how to tune performance, protect data, and design for high availability
  • Experience with storage networking and protocols like NFS, SMB, iSCSI, Fibre Channel, RDMA, and NVMe-oF
  • Practical experience with automation and infrastructure as code using tools such as Terraform, Ansible, or Puppet
  • Strong knowledge of monitoring and observability tools (for example Prometheus, InfluxDB, or Elastic stack), logging, and alerting used to operate and improve storage systems

Desired Qualifications

  • Experience managing and scaling SRE/Production Engineering teams in large scale, mission critical environments with focus on operational excellence, service availability, and performance
  • A track record of improving reliability, simplicity, and day to day operations for large, business critical storage systems
  • Experience building or scaling storage for AI/ML or HPC workloads, including hybrid or multi cloud setups (for example AWS S3, Azure Blob, or Google Cloud Storage), as well as on-prem infrastructure
  • Experience with software defined storage, cloud native storage, and Kubernetes based storage orchestration
  • A visible passion for mentoring, career development, and building a supportive, high performing team culture

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce