Data Platform Engineer
$170,000–$300,000 year
On-siteLos Angeles, California, United States
Job Summary
Build the data backbone for autonomous factories by transforming machine signals, quality events, and work orders into trusted datasets for scheduling and ML systems. Own end-to-end data flows including ingestion, CDC, streaming, and lakehouse storage while managing ordering, retries, and idempotence. Define versioned data contracts with upstream teams and model telemetry using explicit grain, identity, and provenance. Operate Dagster orchestration, dbt transformations, and data CI/CD pipelines to ensure reliability for analytics and operations. Collaborate with Manufacturing Operations to acquire data from PLCs, historians, and industrial protocols via OPC-UA, MTConnect, and MQTT.
Required Qualifications
- Experience building and operating production data infrastructure or distributed data systems, including on-call ownership and recovery efforts
- Strong production Python
- Advanced SQL
- Data-modeling skills, including incremental processing, temporal data, and schema evolution
- Experience with Kafka or another event-streaming platform
- Experience with CDC or other stateful incremental pipelines
- Experience operating Snowflake
- Experience with a lakehouse table format such as Iceberg, Delta, or Hudi, including expertise in partitioning and compaction
- Experience with tools such as Dagster, Airflow, Argo, or Prefect
- Experience with dbt or similar transformation frameworks
- Experience with Kubernetes or infrastructure as code
- Strong judgment regarding contracts, failure modes, and the needs of downstream analytics, ML, and operational systems
Desired Qualifications
- Controls experience
- Experience running Snowflake and Iceberg together or designing a hybrid warehouse and lakehouse architecture
- Production experience with PeerDB, Debezium, Flink, Spark Structured Streaming, Redpanda, Bufstream, or similar CDC and streaming systems
- Proficiency with ClickHouse or another low-latency analytical database, including performance tuning and lifecycle management
- Experience with industrial or edge data collection using OPC-UA, MTConnect, MQTT, historians, PLCs, or handling intermittently connected systems
- Background in performance-sensitive data systems built with Go, Rust, or Scala
- Regulated-environment experience
- Contributions to dbt, Dagster, Iceberg, or related projects
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.