Data Engineer
HybridChantilly, Virginia, United States
Job Summary
Design and maintain ETL/ELT pipelines for batch and real-time processing using Python and SQL to structure large volumes of data for custom enterprise systems. Integrate sources including databases, APIs, streaming platforms, and PDFs while building scalable architectures to support analytics and machine learning workloads. Optimize data processing and queries for performance in AWS S3, implement data governance practices, and develop web scraping workflows to collect open-source datasets. Collaborate with Data Scientists to prepare feature-ready datasets, support ML model deployment, and utilize Docker, Kubernetes, and CI/CD pipelines for workflow management. Work on client site in a fast-paced environment requiring 3–5 years of experience.
Required Qualifications
- 3–5 years + of professional experience in data engineering or related roles
- Strong collaboration skills to work effectively with Data Scientists, Analysts, and Engineering teams
- Ability to communicate complex technical concepts to non-technical stakeholders
- Detail-oriented, curious, and committed to data quality
- Capable of managing multiple priorities in a fast-paced environment
- Python
- SQL
- AWS cloud services
- Linux environments
- Git
Desired Qualifications
- PySpark
- Elastic/OpenSearch
- Understanding of machine learning workflows and MLOps concepts
- Exposure to PySpark or other big data frameworks
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.