Weekday logo
WeekdayPosted 2 weeks ago
EXPIRED

Senior Research Engineer

$5,000,000–$9,000,000 year

On-siteBengaluru, Karnataka, India

Full TimeMid LevelSmall

Job Summary

Design and maintain comprehensive LLM evaluation frameworks to measure model capabilities, reliability, and safety. Conduct research on model behavior, prompting, fine-tuning, and evaluation methodologies while developing novel experiments for emerging capabilities. Create high-quality datasets, test suites, and automated pipelines to analyze outputs and identify performance gaps. Build scalable tooling for running evaluations across large numbers of prompts and models. Collaborate with researchers and product teams to translate insights into measurable improvements, investigating failures to capture undetected weaknesses. Contribute to technical documentation and internal benchmarks.

Required Qualifications

  • 1–8 years of experience in machine learning, AI research, software engineering, data science, or a related technical field.
  • Strong hands-on experience designing and implementing LLM evals or model evaluation systems.
  • Solid understanding of LLM research, including model capabilities, prompting, fine-tuning, inference, and evaluation methodologies.
  • Strong Python programming and experience working with ML/AI frameworks and data-processing pipelines.
  • Ability to formulate research questions, design experiments, interpret results, and communicate technical findings clearly.
  • Strong analytical and problem-solving skills with attention to experimental rigor and reproducibility.
  • Experience working with large datasets, automated testing, and evaluation pipelines.

Desired Qualifications

  • Experience developing or working with LLM benchmarks and standardized evaluation suites.
  • Familiarity with benchmark design, dataset curation, scoring methodologies, and statistical analysis.
  • Experience with open-source LLMs, model APIs, Hugging Face, PyTorch, or similar frameworks.
  • Exposure to reinforcement learning, RLHF/RLAIF, fine-tuning, synthetic data generation, or agentic systems.
  • Research publications, technical blogs, open-source contributions, or demonstrated independent AI research work.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Find similar roles