Synthlane logo
SynthlanePosted 1 week ago

Data Engineer – Mid & Senior Level

RemoteUnited States

Part TimeSenior LevelSmall

Job Summary

Design and maintain production-grade data pipelines on AWS that transform operational data into high-quality, secure, and AI-ready datasets. Extract, transform, validate, and curate large-scale Parquet datasets while implementing data de-identification, masking, and privacy-preserving transformations. Manage pipeline orchestration, scheduling, retries, and backfills using tools like Airflow, Dagster, or AWS Step Functions. Build reliable, idempotent pipelines with comprehensive data quality checks, monitoring, and alerting. Contribute to CI/CD pipelines, Infrastructure as Code practices, and maintain data catalogs, metadata, and lineage. Provide production support, troubleshooting, and root cause analysis for data pipeline issues. Follow software engineering best practices including Git, code reviews, and automated testing.

Required Qualifications

  • Strong proficiency in Python
  • Strong proficiency in SQL
  • Hands-on experience with AWS data services
  • Hands-on experience with production data pipelines
  • Experience with Apache Spark
  • Practical experience with Airflow
  • Practical experience with Dagster
  • Practical experience with AWS Step Functions
  • Strong understanding of data pipeline architecture
  • Strong understanding of ETL/ELT
  • Strong understanding of data transformation
  • Experience working with Parquet
  • Experience working with large-scale datasets
  • Understanding of data quality
  • Understanding of schema management
  • Understanding of monitoring
  • Understanding of alerting
  • Experience with Git
  • Experience with code reviews
  • Experience with automated testing
  • Experience with idempotency
  • Experience with error handling
  • Experience with retries
  • Experience with backfills
  • Experience supporting production data pipelines
  • Ability to work effectively with cross-functional engineering teams
  • Ability to work effectively with data teams
  • Data De-identification
  • Data Lineage
  • Data Catalog
  • CI/CD
  • Data Orchestration
  • Production Support
  • Data Transformation
  • Schema Management
  • Data Quality
  • Data Engineering
  • Data Pipelines
  • ETL/ELT
  • AWS
  • Python
  • SQL
  • Apache Spark
  • Airflow
  • Dagster
  • AWS Step Functions
  • Parquet

Desired Qualifications

  • Debezium
  • AWS DMS
  • Apache Iceberg
  • Delta Lake
  • Apache Hudi
  • Data Masking
  • Data Tokenization
  • Terraform
  • CloudFormation
  • ML/AI Training Data
  • Privacy-Preserving Data

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce