Data Engineer
On-siteToronto, Ontario, Canada
Toronto, Ontario, CanadaOn-siteFull TimeSmall
Full TimeSmall
Job Summary
Build end-to-end ETL pipelines using Azure Data Bricks, PySpark, and Scala, covering batch autoloader jobs and structured streaming. Design and implement data warehousing solutions including SCD1, SCD2, and CDC pipelines while tuning Spark performance and managing Delta Lake partitions. Develop Lakehouse federation architectures to connect foreign catalogs and external data sources, ensuring governance, security, and CI/CD integration. Create Unity Catalog environments, schemas, tables, and materialized views to support scalable data modeling.
Required Qualifications
- Strong hands-on experience with Data-bricks and Apache Spark (PySpark/Scala)
- Experience in SQL and data transformation techniques
- Knowledge of ETL tools and data pipeline development
- Experience working with cloud platforms (Azure/AWS/GCP)
- Strong Azure cloud background
- Understanding of data warehousing concepts
- Strong problem-solving and analytical skills
- Hands-on experience with Azure Data bricks or Delta Lake, in building ETL pipelines : batch (autoloader) and Spark structured streaming
- Knowledge of data modelling and performance tuning in Spark
- Exposure to CI/CD pipelines and DevOps practices
- Familiarity with data governance and security practices
- Strong hands on working experience of Unity catalog
- Hands on exposure to Creating end to end environments : creating catalogs, schemas, tables . materialized views, functions, volumes
- Experience in building SCD 1 and SCD2 (slowly changing dimensions ) on dimension tables
- Experience in building CDC (change data capture pipelines)
- Strong hands on experience with Lakehouse federation , creating foreign catalogs to get data from external sources
- Strong understanding of databricks partitioning , Liquid clustering
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.