Data Engineer
On-siteLondon, England, United Kingdom
Job Summary
Build reusable, scalable data pipelines and modernize the platform using Snowflake, Databricks, and Spark to support over $75 trillion in assets under management. Design financial data models, automate ingestion and quality assurance workflows, and optimize operational costs while exploring agentic AI patterns for pipeline automation. Collaborate with cross-functional teams to implement technical solutions that meet business needs, ensuring compliance with financial data standards. Work with Python, SQL, and cloud infrastructure to deliver production-quality code and share best practices through code reviews.
Required Qualifications
- 3–5 years of hands-on experience in data engineering or a closely related discipline
- Demonstrated experience building shared tooling, frameworks, or reusable components, not only end-to-end pipelines
- Strong proficiency in Python; comfortable writing production-quality, well-tested code
- Advanced SQL for data modelling, query optimisation, and analytical work
- Hands-on, mandatory experience with Databricks and Spark for large-scale distributed data processing, including Delta Lake, Spark SQL, and cluster optimisation
- Experience with Snowflake or equivalent cloud warehouses (BigQuery, Redshift, Synapse) alongside Databricks
- Working knowledge of at least one workflow orchestrator Airflow, Prefect, or Dagster
- Practical experience on AWS, Azure, or GCP object storage, compute, serverless, IAM
- Comfortable with Git, Docker, and CI/CD pipelines for data platform deployments
- Experience implementing data quality checks, schema validation, or contract testing
- Proficient in using AI coding tools such as GitHub Copilot and Claude to accelerate development, generate boilerplate, review code, and navigate complex codebases. Comfortable integrating these tools into a daily engineering workflow
Desired Qualifications
- Experience working with financial or enterprise data environments
- Hands-on experience with agentic AI frameworks (LangChain, LlamaIndex, AutoGen, CrewAI, or the Anthropic Agent SDK)
- Knowledge of streaming data processing (Kafka, Kinesis, or Flink)
- Exposure to financial data types pricing, reference data, benchmarks, indices, or corporate actions
- Experience with metadata catalogues (Unity Catalog, DataHub, OpenMetadata, Alation, or similar)
- Familiarity with data contract patterns
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.