Member of Technical Staff — Data Infrastructure
On-siteSan Francisco, California, United States
Job Summary
Design and operate petabyte-scale storage using lakehouse architecture optimized for batch and real-time queries. Own the shared compute and orchestration platform, including Spark, Ray, and workflow scheduling, for ingestion and research pipelines. Optimize data strategy end-to-end from storage to loading, managing high-throughput data loading up to the tensor boundary. Build systems for cataloging, deduplication, lineage, search, and reproducibility across the data lifecycle. Implement platform-level quality and monitoring tooling, while scaling infrastructure to improve engineering velocity with dedicated alerting. Work across the full data lifecycle, including building and operating ingestion pipelines for critical data sources.
Required Qualifications
- Demonstrated experience building large-scale data pipelines and distributed compute systems (e.g. Spark, Ray, Beam)
- Knowledge of state-of-the-art methods and tools for data ingestion, storage, and loading — including file formats and storage systems (e.g. Parquet, Zarr, Delta Lake) and how they impact performance and scalability
- Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines
- Understanding of how data loading throughput affects large-scale training, and experience optimizing it
- Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution
Desired Qualifications
- We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.