ComplyAdvantage logo
ComplyAdvantagePosted 1 month ago

Principal Data Engineer

$92,000–$115,000 year

HybridLisbon, Lisbon, Portugal

Full TimeSenior LevelMediumFinancial Technology

Job Summary

Set medium-to-long term technical direction for the petabyte-scale data platform powering AML, KYC, and fraud products. Architect complex systems spanning ingestion, transformation, and serving of billions of daily signals into a real-time financial crime knowledge graph. Lead design of foundational infrastructure including event buses, lakehouses, and orchestration pipelines while establishing standards for data quality, observability, and schema evolution. Coach engineers across the organization and represent ComplyAdvantage at industry events. Own architecture decisions for ML, agentic AI, and retrieval pipelines supporting product teams.

Required Qualifications

  • Substantial experience designing and operating production-grade data platforms at high scale, whether that is large request volumes, large data volumes, or both.
  • Deep expertise in distributed data systems: streaming (Kafka or similar), batch and ELT/ETL frameworks (Spark, Flink, dbt, Airflow or Argo Workflows), and modern lakehouse or warehouse technologies.
  • Strong production Python experience and awareness of other relevant languages (Java/Kotlin) sufficient to set direction, review code and coach others.
  • Experience designing for cloud (AWS and GCP) and containerised infrastructure (Kubernetes, Docker, ArgoCD).
  • A track record of treating data quality, observability and data contracts as first-class engineering concerns rather than afterthoughts.
  • Strong working understanding of logging, monitoring, alerting and incident management tooling for data systems.
  • Excellent written and verbal communication. You can produce technical documentation that senior leaders and engineers can act on.
  • Ownership of software and data products from inception through to production and long-term operation.
  • A track record of coaching staff, senior and mid-level engineers, and of helping Recruiting improve hiring and onboarding.
  • Architect petabyte-scale data platforms across batch, micro-batch and streaming, making explicit trade-offs between latency, throughput, cost and operational complexity.
  • Design and own the lineage, quality, freshness and observability of the financial crime knowledge graph and the pipelines that feed it.
  • Build and evolve the foundational data infrastructure: ingestion frameworks, the event bus, feature and serving stores, the lakehouse, orchestration and the developer experience around them.
  • Set the standard for event-sourced and streaming patterns across the company using Kafka and similar technologies, and drive consistency in how services produce and consume data.
  • Design data services with scale and ease of operation in mind. Write maintainable, performant, well-tested Python code (and where appropriate Kotlin or Python), and review the work of others.
  • Partner with ML engineers and data scientists so the platform supports feature engineering, training pipelines and online inference at scale.
  • Set the data quality, schema evolution and contract-testing standards that other engineering teams adopt.
  • Integrate the data platform with new and existing services. Build and consume APIs and event streams, and produce documentation that engineers and analysts can self-serve from.
  • Coach staff, senior and mid-level engineers across the tribe and the wider engineering organisation, and build the bench of future technical leaders.
  • You will own the technical architecture of the data platform behind our sanctions, PEP, adverse media, transaction monitoring, fraud and customer risk products.
  • You will lead the architecture that supports ML, data science, and product teams ship new detection models and risk signals in days rather than quarters.
  • You will design the data foundations that make agentic AI work at scale: retrieval pipelines, grounding sources, tool data and the event histories that let agents reason over our knowledge graph.
  • You will work hand in hand with our Customer Risk, Fraud, Knowledge Graph and Screening tribes so the data foundations keep pace with the product and AI roadmap.
  • You will set the technical direction for how we ingest, normalise and merge entity, relationship and event data from millions of public and private sources.
  • You will be the deciding voice on company-wide data architecture decisions, the make-or-buy choices that follow, and our long-term vendor and tooling strategy for the data estate.
  • Development is organised around Kotlin and Python for our backend languages and TypeScript/ES6+React for our frontend stack
  • We make substantial use of relational database technologies, notably Postgres, Yugabyte
  • We also use an event-sourced model powered by Kafka for our communication bus and gRPC for our intra-service communication protocol
  • We use modern observability solutions from Grafana Cloud and deploy our code using ArgoCD
  • We embrace a hybrid approach that requires employees to be in the office for two days a week.

Desired Qualifications

  • Experience building or operating data systems in financial services, AML, KYC, fraud, regtech or another regulated domain.
  • Familiarity with knowledge graph and entity resolution problems: deduplication, linkage, hierarchies and temporal relationships.
  • Experience supporting ML, LLM and agentic AI workloads, including feature stores, vector stores, retrieval pipelines, tool data and online/offline parity.
  • Experience representing engineering externally at conferences, meet-ups or in technical publications.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce