Yubi logo
YubiPosted 8 months ago

Data Scientist -II

On-siteChennai, Tamil Nadu, India

Full TimeDoctorate Or Professional DegreeMedium

Job Summary

Develop scalable, reusable machine learning tools specializing in computer vision and natural language processing to drive innovative solutions for image classification, data extraction, text classification, and summarization. Collaborate with cross-functional teams to design and implement efficient ML pipelines that facilitate seamless model integration and deployment in production environments. Spearhead the optimization of the model development lifecycle, focusing on scalability for training and production scoring to manage significant data volumes and user traffic. Implement cutting-edge technologies to enhance model training throughput and response times, leveraging advanced frameworks like TensorFlow, Keras, and Fast API. Architect reusable APIs to integrate OCR capabilities across diverse applications, overcoming challenges with libraries such as Tesseract and EasyOCR.

Required Qualifications

  • 3+ years of experience in developing computer vision & NLP models and applications
  • Extensive knowledge and experience in Data Science and Machine Learning techniques, with a proven track record in leading and executing complex projects
  • Deep understanding of the entire ML model development lifecycle, including design, development, training, testing/evaluation, and deployment, with the ability to guide best practices
  • Expertise in writing high-quality, reusable code for various stages of model development, including training, testing, and deployment
  • Advanced proficiency in Python programming, with extensive experience in ML frameworks such as Scikit-learn, TensorFlow, and Keras and API development frameworks such as Django, Fast API
  • Demonstrated success in overcoming OCR challenges using advanced methodologies and libraries like Tesseract, Keras-OCR, EasyOCR, etc
  • Proven experience in architecting reusable APIs to integrate OCR capabilities across diverse applications and use cases
  • Proficiency with public cloud OCR services like AWS Textract, GCP Vision, and Document AI
  • History of integrating OCR solutions into production systems for efficient text extraction from various media, including images and PDFs
  • Comprehensive understanding of convolutional neural networks (CNNs) and hands-on experience with deep learning models, such as YOLO, DETR
  • Strong capability to prototype, evaluate, and implement state-of-the-art ML advancements, particularly in OCR and CV-NLP
  • Extensive experience in NLP tasks, such as Named Entity Recognition (NER), text classification, and on fine tuning of Large Language Models (LLMs)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce