Data Engineer – Mid & Senior Level
RemoteUnited States
Job Summary
Design and maintain production-grade data pipelines on AWS that transform operational data into high-quality, secure, and AI-ready datasets. Extract, transform, validate, and curate large-scale Parquet datasets while implementing data de-identification, masking, and privacy-preserving transformations. Manage pipeline orchestration, scheduling, retries, and backfills using tools like Airflow, Dagster, or AWS Step Functions. Build reliable, idempotent pipelines with comprehensive data quality checks, monitoring, and alerting. Contribute to CI/CD pipelines, Infrastructure as Code practices, and maintain data catalogs, metadata, and lineage. Provide production support, troubleshooting, and root cause analysis for data pipeline issues. Follow software engineering best practices including Git, code reviews, and automated testing.
Required Qualifications
- Strong proficiency in Python
- Strong proficiency in SQL
- Hands-on experience with AWS data services
- Hands-on experience with production data pipelines
- Experience with Apache Spark
- Practical experience with Airflow
- Practical experience with Dagster
- Practical experience with AWS Step Functions
- Strong understanding of data pipeline architecture
- Strong understanding of ETL/ELT
- Strong understanding of data transformation
- Experience working with Parquet
- Experience working with large-scale datasets
- Understanding of data quality
- Understanding of schema management
- Understanding of monitoring
- Understanding of alerting
- Experience with Git
- Experience with code reviews
- Experience with automated testing
- Experience with idempotency
- Experience with error handling
- Experience with retries
- Experience with backfills
- Experience supporting production data pipelines
- Ability to work effectively with cross-functional engineering teams
- Ability to work effectively with data teams
- Data De-identification
- Data Lineage
- Data Catalog
- CI/CD
- Data Orchestration
- Production Support
- Data Transformation
- Schema Management
- Data Quality
- Data Engineering
- Data Pipelines
- ETL/ELT
- AWS
- Python
- SQL
- Apache Spark
- Airflow
- Dagster
- AWS Step Functions
- Parquet
Desired Qualifications
- Debezium
- AWS DMS
- Apache Iceberg
- Delta Lake
- Apache Hudi
- Data Masking
- Data Tokenization
- Terraform
- CloudFormation
- ML/AI Training Data
- Privacy-Preserving Data
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.