Reliability Engineer
$122,440–$232,190 year
On-siteSanta Clara, California, United States or Beaver Brook, Massachusetts, United States
Job Summary
Define and maintain pod-level reliability/availability specs and targets (MTBF, AFR, RAS) for compute, memory, storage, network, power, and cooling subsystems. Translate system/SLA requirements into pod and subsystem level reliability specs, flowing requirements down to silicon, platform, and facilities teams. Lead FMEA, root-cause analysis, and pod fleet failure-data analytics to drive corrective actions and spec updates. Architect RAS features (ECC, memory mirroring, predictive failure, telemetry) and graceful degradation/redundancy against pod-level specs. Partner with facilities on pod power/cooling redundancy (N+1, 2N), thermal margins, and disaster-recovery readiness. Establish HALT/HASS, burn-in, qualification processes; track field returns and KPIs against pod spec.
Required Qualifications
- BS/MS/PhD in EE/ME Reliability or related
- at least 4-6 yrs experience
- Experience authoring and owning reliability specs and requirement flow-down
- Strong RAS skills
- Strong FMEA skills
- Strong statistical reliability (Weibull, FIT) skills
- Experience with large-scale fleet telemetry
- Experience with thermal/power redundancy
- Must be able to work on-site in US, Massachusetts, Beaver Brook
- Must be able to work on-site in US, California, Santa Clara
Desired Qualifications
- AI cluster operations
- data analytics (Python/SQL)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.