ML Engineer
RemoteUnited Kingdom or Serbia
Job Summary
Investigate why documents fail at population scale by querying production datasets, comparing data representations, and finding statistical patterns that explain hundreds of failures. Own and evolve the accuracy measurement framework, ensuring every fix has an expected uplift, measured uplift, and post-deployment monitoring plan. Build an embedding-based entity matching service and an OCR correction layer to fix extraction errors before pipeline processing, evaluating candidates like vision-capable LLMs and alternative OCR engines. Design fixes for pipeline logic errors by reading the C# codebase, understanding processing step sequences, and identifying root causes. Set up ML pipeline orchestration and MLOps practices including experiment tracking, model versioning, and containerised model serving. Collaborate with C# backend engineers, a product manager, and the AI team on model architecture and deployment infrastructure.
Required Qualifications
- UK, Serbia, or Moldova location
- Self-directed and autonomous working style
- Comfortable reading and tracing C# / .NET code
- Investigation & Measurement: Measurement rigour as a core discipline
- Investigation & Measurement: Forensic data investigation at scale
- Investigation & Measurement: Strong SQL (PostgreSQL, complex analytical queries)
- Investigation & Measurement: Strong Python (pandas, NumPy, scikit-learn)
- ML Engineering - Structured Document Processing: Hands-on experience with the structured document processing domain
- ML Engineering - Structured Document Processing: Practical ML skills to deliver projects end-to-end
- ML Engineering - Structured Document Processing: LLM integration for structured data tasks
- ML Engineering - Structured Document Processing: MLOps and deployment
- Collaboration: Collaboration across disciplines
Desired Qualifications
- Financial document or accounting domain knowledge (invoices, charts of accounts, tax treatment, Xero/QBO)
- Experience with managed OCR services (Azure DI, Google Cloud Vision, AWS Textract) or open-source alternatives
- Experience with pre-trained document understanding models (LayoutLM, Donut, or similar)
- Experience building LLM-as-judge or LLM-as-corrector evaluation systems
- Experience with document-oriented processing tools (docling, pdfplumber, PyPDF, or equivalents)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.