Senior Databricks Data Engineer
HybridBucharest, București, Romania
Job Summary
Design and implement scalable ETL/ELT pipelines using PySpark, Scala, and Databricks SQL within a Medallion architecture. Optimize Delta Lake tables and Spark clusters for performance, while building real-time solutions with Structured Streaming and Delta Live Tables. Administer Unity Catalog for governance, enforce data quality rules, and manage complex workflows via Databricks Jobs or external orchestration tools. Integrate pipelines into CI/CD processes using Git, Bundles, and Terraform. Provide technical mentorship to junior developers and collaborate with Data Scientists and Architects on business requirements.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Mathematics, or a relevant technical field
- Minimum of 5+ years of experience in Data Engineering
- At least 3+ years using Databricks and Spark at scale
- Proven expert-level experience with the entire Databricks ecosystem (Workspace, Cluster Management, Notebooks, Databricks SQL)
- In-depth knowledge of Spark architecture (RDD, DataFrames, Spark SQL) and advanced optimization techniques
- Expertise in implementing and managing Delta Lake (ACID properties, Time Travel, Merge, Optimize, Vacuum)
- Advanced/expert-level proficiency in Python (with PySpark) and/or Scala (with Spark)
- Advanced/expert-level skills in SQL and Data Modeling (Dimensional, 3NF, Data Vault)
- Solid experience with a major Cloud platform (AWS, Azure, or GCP), particularly with storage services (S3, ADLS Gen2, GCS) and networking
- Expertise in implementing and optimizing the Medallion architecture (Bronze, Silver, Gold) using Delta Lake
- Efficient implementation of the Lakehouse architecture on Databricks, combining best practices from Data Warehousing (DWH) and Data Lakes
- Designing and implementing real-time/quasi-real-time data processing solutions using Spark Structured Streaming and Delta Live Tables (DLT)
- Implementing and administering Unity Catalog for centralized data governance, fine-grained security (row/column-level security), and data lineage
- Defining and implementing data quality standards and rules (e.g., using DLT or Great Expectations) to maintain data integrity
- Developing and managing complex workflows using Databricks Workflows (Jobs) or external tools (e.g., Azure Data Factory, Airflow) for pipeline automation
- Integrating Databricks pipelines into CI/CD processes using tools like Git, Databricks Repos, and Bundles
Desired Qualifications
- Unity Catalog: Practical experience with implementing and managing Unity Catalog
- Experience with Delta Live Tables (DLT) and Databricks Workflows
- Understanding of basic MLOps concepts and experience with MLflow to facilitate integration with Data Science teams
- Experience with Terraform or equivalent for Infrastructure as Code (IaC)
- Databricks certifications (e.g., Databricks Certified Data Engineer Professional) are a significant advantage
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.