3 months ago
Base Salary
$293k - $325k/yr
Responsibilities
- Build web applications, APIs, data models, and backend services for AI research workflows.
- Develop tools for authoring and managing evaluation tasks, rubrics, graders, suites, and rollout configurations.
- Create publishing, versioning, auditing, and sharing workflows for research artifacts.
- Automate evaluation runs and generate reports for design, research, and engineering teams.
- Support synthetic data generation workflows involving transcripts, media, and model comparisons.
- Translate research and product questions into measurable scenarios, automated graders, and human-evaluation campaigns.
- Develop measures of task quality, coverage, diversity, and semantic spread.
- Diagnose issues across application code, workers, model endpoints, deployments, and compute infrastructure.
- Improve reliability through health checks, observability, reproducible launch paths, data integrity safeguards, and automated verification.
- Lead migrations and dependent changes across research tools, evaluation systems, and supporting services.
- Collaborate with designers, model researchers, research engineers, and infrastructure teams, and onboard contributors.
Requirements
- 7+ years of professional software engineering experience.
- Strong full-stack experience across web applications, backend services, APIs, and data models.
- Experience owning complex systems spanning multiple services or repositories.
- Expertise in generative AI, multimodal models, or model-evaluation systems.
- Experience building internal tools for technical and non-technical users.
- Ability to debug distributed workflows and production infrastructure.
- Strong product judgment and ability to turn ambiguous requirements into concrete plans.
- Clear communication and effective collaboration across engineering, design, and research.
- Preferred experience with synthetic data generation, simulation, conversational AI, speech, video, motion, or embodied interaction.
- Preferred experience with automated graders, human evaluation, supervised fine-tuning, reinforcement learning, or experiment- and dataset-management platforms.
- Preferred experience operating GPU-backed inference or rollout workloads at very large scale.
Benefits
- Based in San Francisco, California.
- Hybrid work model with four days in the office per week.
- Relocation assistance is offered to new employees.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
