Linguist II
RemoteWashington, United States
Job Summary
Apply linguistic expertise in syntax, semantics, pragmatics, and sociolinguistics to support large language models and voice-enabled technologies. Collaborate with linguists, data operations teams, and ML engineers on data collection, curation, annotation, and localization efforts for model training and fine-tuning. Contribute to the development of annotation schemas and guidelines for instruction-tuning, preference labeling, and RLHF. Evaluate datasets used for pre-training, fine-tuning, and alignment, while supporting programmatic methods for generating synthetic annotated data at scale. Assist in model evaluation efforts including prompt-based testing, red-teaming, and linguistic error analysis. Contribute to AI safety and responsible AI practices through linguistic review of model outputs for hallucination detection and bias identification. Participate in experiments to assess data quality, annotation consistency, and downstream model performance.
Required Qualifications
- Bachelor's degree in Linguistics, Computational Linguistics, Computer Science, Speech Science, or related field
- 1+ years of experience in Linguistics, Language Technologies, NLP, or AI/ML data operations (or equivalent)
- Native or near-native fluency in English and at least one additional language
- Knowledge of syntax, semantics, pragmatics, sociolinguistics, corpus linguistics, and other areas of linguistics
- Familiarity with Large Language Models (LLMs), their applications and data practices (training data, evaluation, prompting, fine-tuning)
- Exposure to LLM evaluation methodologies (human evaluation, automated metrics, adversarial testing)
- Experience working with semantic ontologies, taxonomies, or intent/slot frameworks
- Proficiency using AI Agents/Chatbots
- Experience with database queries and data analysis processes (SQL, spreadsheets, R, Unix, or others)
- Experience working with speech and text data in multiple languages
- Comfortable working in a fast-paced, highly collaborative environment with evolving priorities
Desired Qualifications
- Master's degree in Linguistics, Computational Linguistics, Language Technologies, or a related field
- Familiarity with machine learning frameworks, NLP libraries, and tools (e.g., Hugging Face, spaCy, NLTK, PyTorch)
- Exposure to statistical language modeling or training data pipelines
- Strong organizational skills and attention to detail
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.