Base Salary
$230k - $385k/yr
Responsibilities
- Own the production improvement loop across agent behavior, customer and operator feedback, evaluation, experimentation, and verified business outcomes.
- Instrument agent workflows to make model interactions, tool use, decisions, failures, human edits, and downstream outcomes observable.
- Define quality standards, representative evaluation datasets, regression coverage, and production monitoring for GTM workflows.
- Investigate agent underperformance across context, knowledge, instructions, tools, routing, guardrails, and workflow design.
- Design and ship behavior improvements involving prompting, context construction, decision logic, tool use, and human-review paths.
- Build backend services, APIs, data models, and feedback pipelines that make agent behavior observable, steerable, and reproducible.
- Run controlled experiments, production replays, and staged rollouts to measure quality and business outcomes.
- Partner with Engineering, Product, Data Science, Sales, and B2B Marketing to prioritize high-value problems and define success.
- Ship safeguards for privacy, security, reliability, human oversight, and safe operational rollout.
Requirements
- At least 4 years of experience in software engineering, backend engineering, applied AI, or product engineering building reliable production systems.
- Experience building AI agents, LLM-powered applications, or other model-driven workflows operating on real production traffic.
- Experience diagnosing and improving agent behavior using production traces, user feedback, evaluation, experimentation, or systems design.
- Practical experience with evaluation design, regression testing, human or model grading, online quality signals, or controlled experiments.
- Strong backend engineering skills with Python, APIs, data pipelines, stateful workflows, and production services.
- Ability to connect technical changes to customer experience, conversion, qualified pipeline, or operational efficiency.
- Comfort working across model behavior, context, knowledge, tools, workflow state, and human-in-the-loop decisions.
- Ability to collaborate with technical and non-technical partners across Engineering, Product, Data Science, Sales, and B2B Marketing.
- Experience with agent evaluation, observability, experimentation, or AI infrastructure products is a plus.
- Experience with production replay, LLM grading, human-labeled datasets, shadow evaluation, or staged rollout is a plus.
- Experience improving model or agent behavior through context design, prompting, tools, decision logic, or feedback loops is a plus.
- Experience with sales, B2B marketing, revenue, CRM, campaign, or other GTM-facing systems is a plus.
Tech Stack
Categories
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. AI is an extremely powerful tool that must be created with safety and human needs at its core. OpenAI is dedicated to putting that alignment of interests first — ahead of profit. To achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Our investment in diversity, equity, and inclusion is ongoing, executed through a wide range of initiatives, and championed and supported by leadership. At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.
