Data Scientist - NLP
RemoteUnited States
Job Summary
Collect, clean, and prepare datasets for computational modeling using Python, applying pre-processing functions like stop word removal and tokenization. Engineer NLP features using TF-IDF, word2vec, and GloVe to identify key determinants, then select classification techniques including machine learning, regression, and neural networks to fit business problems. Validate model results, justify findings, and visualize insights to address organizational challenges. Work remotely on long-term federal client engagements, utilizing transformer architectures and GenAI while maintaining a public trust security clearance.
Required Qualifications
- Master's degree
- PhD preferred in Statistics, Mathematics, Computer Science, or similar
- High degree of experience utilizing SAS, R, or Python to support NLP use cases such as Document Summarization, Named Entity Recognition, Sentiment Analysis, and/or Topic Modeling
- At least four years of experience developing scalable, production-ready NLP solutions using sci-kit learn, Keras, TensorFlow, PyTorch, Spark NLP
- Experience using git/github to version control source code
- Experience leveraging transformer architecture to develop NLP models
- Experience with open source NLP packages such as Gensim, SpaCy, or NLTK
- Experience with BERT, GPT-J, RoBERTa, T5 or other transformers
- Experience with machine translation and transcription of foreign language documents using Microsoft Azure translation services
- Experience working in an AWS cloud environment and with related AWS services such as Bedrock and Textract
- Experience coordinating and maintaining user stories
- Must be a US citizen
- Must be able to obtain and maintain a Public trust security clearance
Desired Qualifications
- Experience with GenAI and Prompt Engineering
- Experience in Databricks and MLFlow
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.