Causal Labs logo
Causal LabsPosted 1 month ago

Member of Technical Staff — Data Infrastructure

On-siteSan Francisco, California, United States

Full TimeSenior LevelStartup

Job Summary

Design and operate petabyte-scale storage using lakehouse architecture optimized for batch and real-time queries. Own the shared compute and orchestration platform, including Spark, Ray, and workflow scheduling, for ingestion and research pipelines. Optimize data strategy end-to-end from storage to loading, managing high-throughput data loading up to the tensor boundary. Build systems for cataloging, deduplication, lineage, search, and reproducibility across the data lifecycle. Implement platform-level quality and monitoring tooling, while scaling infrastructure to improve engineering velocity with dedicated alerting. Work across the full data lifecycle, including building and operating ingestion pipelines for critical data sources.

Required Qualifications

  • Demonstrated experience building large-scale data pipelines and distributed compute systems (e.g. Spark, Ray, Beam)
  • Knowledge of state-of-the-art methods and tools for data ingestion, storage, and loading — including file formats and storage systems (e.g. Parquet, Zarr, Delta Lake) and how they impact performance and scalability
  • Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines
  • Understanding of how data loading throughput affects large-scale training, and experience optimizing it
  • Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution

Desired Qualifications

  • We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce