Data Engineer
$540,000–$540,000 year
On-siteArlington, Virginia, United States
Job Summary
Design, build, and maintain Databricks-based data pipelines and lakehouse capabilities that enable secure data integration, analytics, AI/ML, and operational workloads at enterprise scale. Develop production-grade data-processing solutions using Python, SQL, PySpark, Apache Spark, and Delta Lake to ingest, transform, and deliver mission-critical data. Optimize Spark workloads and manage compute resources while implementing data-quality checks, automated testing, and lineage capabilities. Collaborate with engineers, analysts, and cybersecurity teams to deliver reusable data products and support national defense missions. Requires U.S. citizenship, active DoW Secret clearance, and 4+ years of experience with Databricks, AWS/Azure/GCP, and lakehouse architectures.
Required Qualifications
- U.S. Citizens with an active DoW Secret (or higher) clearance
- Bachelor's degree in Computer Science, Engineering, or a related technical field
- 4+ years of relevant data engineering or software engineering experience
- Hands-on experience developing and operating production data pipelines using Databricks
- Proficiency with Python, SQL, PySpark, Apache Spark, and Delta Lake
- Experience building automated ETL/ELT pipelines for large-scale datasets
- Experience designing and maintaining data models, schemas, tables, and lakehouse architectures
- Experience managing Databricks notebooks, jobs, workflows, and compute resources
- Experience implementing data quality, automated testing, monitoring, lineage, or metadata-management capabilities
- Experience working with Databricks and cloud-based data services in AWS, Azure, or Google Cloud
- Experience working with structured, semi-structured, and unstructured data
- Understanding of lakehouse architecture, data governance, security, privacy, and access-control principles
- Ability to troubleshoot data pipelines, Spark workloads, infrastructure, and applications
Desired Qualifications
- Databricks certification or equivalent demonstrated platform expertise
- Experience supporting DoW, federal, Advana, or other enterprise data environments
- Experience with CI/CD, infrastructure as code, automated testing, and source control
- Experience using Unity Catalog for data governance, lineage, and access control
- Experience developing streaming pipelines with Spark Structured Streaming, Kafka, Kinesis, or Pulsar
- Experience with orchestration tools such as Airflow, Dagster, or Argo Workflows
- Experience with Docker, Kubernetes, or other containerization and orchestration technologies
- Experience building cloud-native data platforms in secure, regulated, classified, or mission-critical environments
- Currently holds, or is willing to obtain within 30 days of employment, an approved certification such as Cloud+, GSEC, Security+, or SSCP
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.