Member of Technical Staff — Data Ingestion & Quality
On-siteSan Francisco, California, United States
Job Summary
Own every dataset end to end, from sourcing new multimodal physical data modalities like sparse sensors and point clouds to securing access via partnerships and vendors. Build petabyte-scale data pipelines using Apache Spark to ingest sources into standardized, training-ready storage, covering both batch and streaming workflows with orchestration and monitoring. Develop quality metrics to measure coverage and correctness, catching subtle inconsistencies such as sensor bias and drift, while implementing automated QA checks to continuously monitor data integrity. Write technical requirements for external data vendors and collaborate with researchers to validate that improved datasets translate into model performance.
Required Qualifications
- Demonstrated experience building large-scale data pipelines, QA systems, or evaluation workflows (e.g. Spark, Ray, Beam)
- Detail-oriented in identifying subtle data inconsistencies and issues that could affect quality, with the ability to understand how quality impacts model performance
- Comfortable going deep on unfamiliar source material — reading format specifications, sensor documentation, and vendor manuals to get ingestion exactly right
- Experience working with external data vendors and partners, from technical evaluation to ongoing feedback
- Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.