5 days ago
Base Salary
$295k - $445k/yr
Responsibilities
- Design evaluations for research judgment, hypothesis generation and testing, and long-horizon experiment execution.
- Turn research workflows and model failures into data and evaluation flywheels.
- Improve model research capabilities through agent harnesses, synthetic data, reinforcement-learning environments, and model training.
- Build and maintain safe, reliable integrations between models and OpenAI’s research infrastructure.
- Develop research agents, experiment-orchestration systems, and sandboxed runtimes for research workflows.
- Create metrics and economic models measuring effects on research productivity, model capabilities, and the safety of internal deployments.
Requirements
- Research or engineering experience across LLM training, model evaluations, agent systems, synthetic data, research infrastructure, or large-scale distributed systems.
- Ability to work as a strong generalist across open-ended research and practical implementation.
- Ability to collaborate across systems, data, model training, evaluations, and research teams.
- Comfort building and maintaining data pipelines, tooling, and infrastructure for emerging AI capabilities.
- Ability to work effectively on ambiguous problems without established playbooks.
- Rigor regarding scientific quality, research judgment, safety, privacy, reliability, performance, and scale.
- Interest in using increasingly capable AI systems to accelerate meaningful research.
Benefits
- The role is based in San Francisco, California.
- OpenAI provides reasonable accommodations for applicants with disabilities.