Sr. Engineer, Cloud HPC Platform
$120,000–$150,000 year
On-siteSan Jose, California, United States
San Jose, California, United StatesOn-siteFull Time$120,000–$150,000 yearSenior LevelStartup
Full TimeSenior LevelStartup
Job Summary
Lead the AWS migration for silicon engineering compute workloads, including EDA, simulation, and verification workflows. Inventory dependencies, define cutover plans, and operate the cloud HPC platform using Slurm, AWS ParallelCluster, and Terraform. Own scheduling, job execution, and self-service virtual desktops while managing licenses, storage, and cost controls. Partner with design, verification, and security teams to validate performance, ensure data correctness, and maintain production reliability through incident response and capacity planning.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent practical experience
- 7+ years building and operating Linux infrastructure, including 3+ years in AWS or a comparable cloud environment
- Deep hands-on experience with Enterprise Linux in production as a System Administrator
- Experience designing or operating HPC, batch compute, large-scale simulation, or similarly compute-intensive platforms
- Deep production experience administering Slurm as the primary HPC scheduler, including partitions, QoS, priorities, fair-share, reservations, preemption, accounting, job arrays and dependencies, failure recovery, upgrades, and integration with AWS ParallelCluster or equivalent cloud capacity
- Strong AWS experience across EC2, IAM, VPC, S3, CloudWatch, Systems Manager, KMS, and high-performance storage services
- Strong infrastructure-as-code skills using Terraform or OpenTofu, including reusable modules, state management, review, and testing
- Experience automating Linux images and configuration with Packer, Ansible, Python, and/or Bash
- Strong knowledge of high-performance and shared storage, Linux file systems, data transfer, backup, and recovery
- Production experience operating secure virtual desktop infrastructure (VDI/DVI) for engineering workloads, preferably Citrix Virtual Apps and Desktops, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policy, profile and storage integration, GPU/graphics, clipboard and file-transfer controls, monitoring, capacity and high availability, and performance troubleshooting
- Strong knowledge of cloud networking, DNS, routing, firewalls, VPN, and hybrid connectivity
- Experience supporting FlexNet/FlexLM or another network-license system
- Experience establishing monitoring, alerting, incident response, capacity management, and cost controls for production infrastructure
- Ability to partner directly with engineers, translate run manifests and workload forecasts into CPU/core, memory, GPU, wall-time, storage I/O and capacity, network, license-token, Slurm partition/reservation, and budget requirements, and communicate capacity and migration risks clearly
- Clear written documentation, design proposals, operating procedures, and post-incident reviews
Desired Qualifications
- Experience supporting semiconductor EDA environments, including Cadence, Synopsys, Ansys, Siemens EDA, PDKs, and IP libraries
- Experience migrating EDA, HPC, simulation, or verification workloads from on-prem infrastructure to AWS
- Experience with Amazon FSx for Lustre, FSx for OpenZFS, EFA, AWS Batch, ParallelCluster, or equivalent HPC services
- Experience designing and operating Citrix Virtual Apps and Desktops or comparable virtual desktop infrastructure (VDI/DVI) for engineering workloads, including application publishing, image and patch lifecycle, SSO/MFA, session brokering and policies, GPU/graphics, profile and storage integration, high availability, monitoring, and performance troubleshooting
- Experience benchmarking workload runtime, queue time, storage I/O, scaling efficiency, quality-of-results, and cost
- Familiarity with GitLab CI, artifact repositories, observability platforms, and controlled release processes
- AWS Professional or Specialty certification
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.