Senior Data Scientist, Classification & Discovery
RemotePoland or India
Job Summary
Design and ship entity resolution and deduplication models directly into production pipelines, including confidence scoring to distinguish reliable matches from ambiguous ones. Own the full lifecycle of these models, managing deployment, monitoring, retraining triggers, and incident response when performance degrades. Build and maintain model-serving infrastructure with precision, recall tracking, confidence calibration, and drift detection running against live traffic. Write production-grade code within streaming and graph-based systems like Kafka and Memgraph to integrate models into real-time pipelines. Partner with security, alerting, and observability teams to apply statistical rigor to broader platform problems such as anomaly detection, alert quality, and telemetry pattern recognition. Mentor analysts on best practices for running models reliably in production while communicating findings to technical and non-technical stakeholders.
Required Qualifications
- Strong applied statistics and machine learning background
- Hands-on experience deploying and operating models in production
- Strong production Python skills
- Direct, hands-on experience with Kafka and streaming or event-driven systems
- Direct, hands-on experience with a graph database such as Memgraph or Neo4j
- Working knowledge of networking fundamentals and common telemetry protocols such as SNMP, syslog, and OpenTelemetry
- Experience with entity resolution, record linkage, or deduplication techniques
- Experience building monitoring and evaluation systems for models already in production
- Comfort working with large-scale telemetry, log, or event data
- Strong communication skills
- Track record of owning models through their full production lifecycle
Desired Qualifications
- Experience with MLOps tooling for model versioning, deployment automation, or CI/CD pipelines for machine learning
- Familiarity with providing structured context to LLMs for reasoning over topology, troubleshooting, or remediation workflows
- Experience with anomaly detection or forecasting applied to operational or monitoring data
- Background in cybersecurity, network detection and response, or infrastructure observability products
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.