Agent Post-Training, Context Research
$295,000–$445,000 year
On-siteSan Francisco, California, United States
Job Summary
Context Researcher on OpenAI’s Agent Post-Training team who will scale compute spent on context and work within the frontier training stack to enable the next paradigm of model training with a clear product interface. Design and run experiments that improve scaling of compute on context; own end-to-end improvements to the post-training stack (RL, data pipelines, graders, reward signals, evals, diagnostics, and model-behavior analysis); build evals and environments to surface future failures; partner with Codex and ChatGPT product teams to translate product signal into model improvements; contribute to early-training and alignment interventions and help decide which integrations and fixes are ready for inclusion in major model runs; drive improvements in training machinery for velocity, reliability, observability, reproducibility, cost, latency, and production readiness; collaborate across research, product, infrastructure, data, evals, and safety boundaries; and shape tools and processes to make agents genuinely useful for developers, enterprises, researchers, and everyday users.
Required Qualifications
- Strong technical fundamentals in machine learning, software engineering, systems, statistics, or a related field.
- Hands-on experience with large language models (LLMs), RL, RLHF/RLAIF, post-training, evals, graders, synthetic data, model training, coding agents, tool-using agents, or production ML systems.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.