Principal Data Engineer
On-siteCape Town, Western Cape, South Africa
Job Summary
Design, ingest, and maintain complex data pipelines across Azure and Snowflake platforms, building robust ELT/ETL products for batch, micro-batch, and near real-time data flows. Assemble datasets meeting functional requirements while optimizing warehouse performance and managing database change control as code. Guide senior and junior team members through technical challenges, demo solutions to stakeholders, and create comprehensive documentation and training materials. Collaborate with Solutions Architects and cross-functional engineering teams to drive data modernization and elevate overall data capability within the fintech ecosystem.
Required Qualifications
- Strong hands-on experience with Snowflake, including warehouse/resource management, semi-structured data handling (VARIANT, JSON), and performance/cost optimisation.
- Experience with dbt (Core or Cloud) - models, tests, macros, snapshots, and documentation generation.
- Experience using schemachange, Flyway, Liquibase, or a similar tool to manage database change control as code.
- Proficient in Jinja templating for building dynamic, reusable SQL/config.
- Advanced working knowledge of SQL (DDL, DML, JSON, XML) and extensive experience managing incremental/batch loading methodologies (CDC, CT, CDC-style watermarking).
- Proven experience building ingestion pipelines from diverse source types: APIs, flat files, relational/NoSQL databases, Azure Table Storage, and web scraping.
- Skilled and experienced in the Azure (or AWS/GCP) cloud platform.
- Advanced understanding of relational data structures, including keys, constraints, and triggers.
- Experience with performance tuning and optimisation of RDBMS and/or cloud data warehouses.
- Experience with relational and NoSQL database technologies (MS SQL Server, MongoDB, CosmosDB, etc.).
- Ability to design and implement conceptual, logical and physical data models that support organisational needs.
- Solid understanding and experience in data modeling, data management and governance methodologies.
- Good understanding of data-related frameworks, methodologies, and patterns.
- Strong analytic skills working with structured, semi-structured and unstructured data sets.
- Proficiency in Python (preferred), Java, or Scala.
- Practical experience applying generative AI / LLM tools to engineering or data workflows (e.g. Copilot, ChatGPT, Claude, or similar).
- Experience implementing CI/CD pipelines through technologies such as GitLab, Azure DevOps, etc.
- Experience deploying data systems in an Infrastructure-as-Code (IaC) manner, preferably using Terraform.
- Strong ability to produce high-quality technical documentation as a routine part of delivery.
- Experience supporting and working with cross-functional teams in a dynamic environment.
- Communicates effectively with both technical and non-technical stakeholders.
Desired Qualifications
- Proficiency in Python (preferred), Java, or Scala.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.