NVIDIA logo
NVIDIAPosted 2 weeks ago

Senior HPC DevOps Engineer, NCS

On-siteTel Aviv, Tel Aviv, Israel

Full TimeSenior LevelEnterprise

Job Summary

Design, implement, and maintain large-scale HPC/AI clusters with state-of-the-art monitoring, logging, and alerting systems. Utilize Infrastructure as Code tools to ensure scalable deployments and develop continuous integration and delivery pipelines to automate processes. Create automation scripts for configuration management, operational monitoring, and complex networking tasks. Perform comprehensive troubleshooting from bare metal to application levels and serve as a technical resource to share best practices with internal teams. Support R&D activities by engaging in proof of concepts and value demonstrations for future improvements.

Required Qualifications

  • B.Sc. in Computer Science, Engineering, or a related field
  • 5+ years of experience
  • Advanced proficiency in programming and scripting languages, with a solid understanding of object-oriented programming principles
  • Familiarity with Jenkins, Ansible, Puppet/Chef
  • Deep understanding of Kubernetes and container-related microservice technologies
  • Hands-on experience with event streaming or message queue technologies e.g. Apache Kafka
  • Experience with multiple storage solutions like Lustre, GPFS, ZFS, and XFS
  • Expertise with virtual systems (VMware, Hyper-V, KVM, Citrix)
  • Familiarity with cloud platforms (AWS, Azure, Google Cloud)

Desired Qualifications

  • Proven networking experience or strong knowledge through professional networking training
  • Architectural Insight: Knowledge of CPU and/or GPU architecture
  • Experience with job scheduling workloads and orchestration tools such as Slurm and Kubernetes

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce