Senior Data Scientist — Transaction Intelligence
On-siteRiyadh, Riyadh Region, Saudi Arabia
Job Summary
Categorize short, noisy merchant strings in mixed English and Latinized Arabic, then resolve them to real merchant entities. Build and maintain end-to-end enrichment models, including evaluation, production code, and confidence-based routing for uncertain cases. Develop pattern recognition, anomaly detection, and behavioral analysis models on transaction streams while backtesting risk scores against realized user behavior. Construct efficient batch pipelines at hundreds-of-millions-record scale and design labeled datasets with regression test suites for every release. Collaborate with product and engineering teams to ship model-backed features within regulated fintech guardrails covering data governance, privacy, and cybersecurity.
Required Qualifications
- 3+ years of applied ML/data science, with models shipped and maintained in production at scale — and knowledge of their failure modes.
- Strong analytical range beyond modeling: exploratory analysis, statistical rigor, and feature design on behavioral/tabular data (SQL fluency assumed).
- Hands-on experience with text similarity, fuzzy matching, or entity resolution on noisy real-world strings.
- Strong Python and the scientific stack (scikit-learn, scipy, numpy), with performance and cost awareness at scale — vectorization, sparse data structures, efficient batch computation.
- Demonstrated maturity working under externally imposed constraints — data governance, privacy, cybersecurity, compliance, infrastructure policy — including collaborating with the teams who own them.
- A track record of turning ambiguous 'the model feels wrong' complaints into measured, regression-tested properties.
- Comfort reading and debugging model code you didn't write, and working directly with product teams on loosely-defined problems.
Desired Qualifications
- Arabic / Arabizi text processing — a strong plus; our data is bilingual with unstable romanization.
- Credit or behavioral risk modeling and validation — scorecards, discrimination and calibration measurement, backtesting against realized outcomes.
- Text embeddings and approximate nearest-neighbor retrieval in resource-conscious settings.
- LLM-assisted labeling, distillation, or weak-supervision pipelines.
- ML lifecycle and batch-serving tooling (experiment tracking, data versioning, distributed task queues).
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.