Irth logo
IrthPosted 2 weeks ago

Data Science ML/GenAI Engineer

$125,505–$125,505 year

RemoteCanada

Full TimeSmall

Job Summary

Design, prototype, and evaluate LLM/NLP solutions for text classification, entity extraction, semantic search, and summarization to address business needs. Translate successful prototypes into production-ready, modular code using object-oriented design principles while analyzing existing codebases and data pipelines. Explore, transform, and analyze data independently using Databricks to support machine learning models and data-processing pipelines. Collaborate with data engineering and application development teams to ensure effective integration between Databricks, data pipelines, models, and application services. Participate in technical discussions to establish development practices that improve the reliability and maintainability of data science solutions.

Required Qualifications

  • 3–5 years of experience in data science, machine learning, or ML engineering, with demonstrated experience developing production-quality code beyond notebooks
  • Master's degree or equivalent professional experience in data science, machine learning, computer science, statistics, or a related field, with a strong foundation in statistics and modeling methodologies
  • Hands-on experience developing NLP and/or LLM-based solutions, including: Retrieval-Augmented Generation (RAG), Embeddings and similarity search, including cosine similarity, Hybrid search combining BM25 and semantic search, Text deduplication techniques such as fuzzy matching and MinHash, Text classification and Named Entity Recognition (NER), Search re-ranking, Prompt engineering
  • Strong proficiency in Python, including experience with pandas, scikit-learn, PyTorch, and other relevant machine learning libraries
  • Strong understanding of object-oriented programming, clean code, and software design principles
  • Strong data analysis skills with experience applying statistical methods such as exploratory data analysis, hypothesis testing, and statistical modeling
  • Hands-on experience with Databricks or a comparable data/analytics platform, including the ability to independently explore, transform, and prepare data
  • Solid understanding of the machine learning lifecycle, including model training, evaluation, deployment, and production monitoring for issues such as model drift and data quality
  • Experience designing and developing REST APIs to expose data, models, or analytical results to downstream applications
  • Strong written and verbal communication skills, with the ability to collaborate effectively across technical teams with different responsibilities and priorities

Desired Qualifications

  • Experience working with news, social media, or other large-scale unstructured text data, including media monitoring, web scraping, or APIs
  • Domain knowledge or professional experience in communications, public relations, media, or related industries
  • Familiarity with the broader LLM ecosystem, including vector databases/stores, RAG frameworks, and fine-tuning NLP models
  • Experience with MLflow or similar experiment-tracking and model-management tools
  • Basic knowledge of TypeScript and Node.js
  • Experience with CI/CD practices and version control, including Git, GitHub Actions, or equivalent tools

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce