Opaque Code logo
Opaque CodePosted 1 month ago

Data Scientist - SME (22759)

On-siteMcLean, Virginia, United States

Full TimeDoctorate Or Professional Degree

Job Summary

Data Scientist to analyze large amounts of data (raw, text, sponsor, etc.) to provide business insight. Responsibilities include applying NLP techniques using Python libraries (Spacy, Gensim, NLTK) and deep learning frameworks (PyTorch, TensorFlow, Keras), leveraging HuggingFace Transformers for NLP tasks, performing text classification and topic modeling, and utilizing encoder-decoder/generative language models. The role involves communicating methodological choices and model results, writing Python scripts to pull data from APIs and relational databases, working with SQL (including CTEs and subqueries), and using tools like GitHub/Jenkins for version control and collaboration. Practical measures of performance, advanced statistical analyses on personnel, intelligence, and performance metrics, and presenting findings to senior leadership are expected. Additional duties include cloud computing development, front-end web work (Flask), semantic search applications, tuning LLMs on custom data, and producing visualizations/dashboards with Tableau. The position requires U.S. citizenship and an active TS/SCI clearance with polygraph, and is located in McLean, Virginia.

Required Qualifications

  • Must be a U.S. Citizen
  • Must have an Active TS/SCI clearance with Polygraph
  • Demonstrated professional or academic experience performing NLP tasks and selecting appropriate Python libraries
  • Demonstrated professional or academic experience with Python NLP packages such as Spacy, Gensim, or NLTK
  • Demonstrated professional or academic experience with deep learning frameworks such as PyTorch, TensorFlow, or Keras
  • Demonstrated professional or academic experience with HuggingFace Transformers library and hub
  • Demonstrated professional or academic experience with text classification and topic modeling in Python (scikit-learn or DL models)
  • Demonstrated academic or professional experience using encoder-decoder and generative language models for NLP tasks
  • Demonstrated experience communicating methodological choices and model results
  • Demonstrated professional or academic experience with SQL (CTEs, set operations, aggregations, nested subqueries)
  • Demonstrated professional or academic experience with version control systems such as GitHub and Jenkins
  • Demonstrated experience leveraging GPUs for accelerated computing
  • Develop practical approaches for measuring performance and presenting data insights to senior leadership
  • Conduct advanced statistical analysis on personnel, intelligence, and performance metrics
  • Assist in selection or development of methodology to conduct research
  • Experience writing Python scripts to pull data from web-based APIs and relational databases
  • Experience with cloud computing development and architecture
  • Experience with front-end web development frameworks such as Flask
  • Experience developing applications for semantic search and tuning LLMs on custom data sets
  • Proficiency with Tableau for visualizations and dashboards

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce