Principal Engineer, Data & Compute
$370,300–$418,200 year
HybridSunnyvale, California, United States
Job Summary
Design and evolve the architecture for allocating and orchestrating training and inference workloads across thousands of GPUs and multiple data centers to ensure optimal throughput, resiliency, and cost efficiency. Build systems enabling fast, reliable access to high-volume sensor and simulation data across geographies while preparing for exabyte-scale operations. Develop foundations for large-scale AI workloads to run seamlessly across hybrid and multi-cloud environments. Act as a trusted partner to leadership in aligning compute investments and architecture with company strategy and performance goals. Uplift the broader engineering org through architectural coaching, technical deep dives, and cultivating a culture of operational excellence. This full-time, in-office role is based in Sunnyvale, CA with a hybrid policy, offering a salary range of $370,300 to $418,200 plus equity.
Required Qualifications
- 10+ years designing and building large-scale distributed systems
- at least 4 years focused on GPU-based cloud infrastructure
- Proven experience enabling large-scale AI training, inference, or computer vision workloads in GPU clusters
- Deep understanding of petabyte-scale data architecture, including storage federation, high-throughput access, and data locality for AI workloads
- Strong technical leadership with a track record of defining and communicating architectural strategy, balancing long-term vision with delivery needs
- A natural mentor with a history of developing engineers and influencing technical direction across teams
- Advanced degree in Computer Science, Electrical Engineering, or a related field
- equivalent industry experience (to Advanced degree)
- This is a full-time role based in Sunnyvale, CA (hybrid)
- This role is a full-time role based in Sunnyvale, CA (hybrid)
Desired Qualifications
- Experience with multi-cloud orchestration, particularly in latency- or cost-sensitive training and inference pipelines
- Familiarity with systems like Ray, Kubernetes, Airflow, or Flyte
- deep fluency in AI/ML job scheduling, model lifecycle management, and infrastructure-as-code practices
- Background in supporting safety-critical or real-time inference use cases (e.g., robotics, autonomous vehicles, aerospace)
- Passion for building infrastructure-as-a-product that delivers performance and simplicity to research and product teams alike
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.