Base Salary
$230k - $385k/yr
Responsibilities
- Design and build the core agent harness and execution loop for interpreting model outputs, using tools, executing code, and completing long-horizon tasks safely.
- Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents in real development environments.
- Develop evaluation, experimentation, and debugging systems that separate harness, model, inference/runtime, and product failures.
- Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to improve solve rate, reliability, latency, and cost.
- Improve observability, profiling, and diagnostics across backend systems, inference, GPUs, and fleet capacity.
- Collaborate with research to make the harness trainable, measurable, and useful for improving frontier agentic models.
- Build shared primitives that make Codex faster, safer, more reliable, and easier for teams and open-source users to extend.
Requirements
- Production experience building or operating systems in distributed systems, infrastructure, developer tooling, sandboxing, virtualization, cloud platforms, or ML systems.
- Hands-on experience with LLM applications, coding agents, evaluation systems, model deployment, inference, compiler or runtime performance, or developer platforms.
- Ability to work across Rust systems code, Python configuration layers, APIs, agent orchestration, evaluations, logs and traces, inference behavior, runtime constraints, and user outcomes.
- Strong focus on reliability, safety, performance, debuggability, and clean abstractions.
- Ability to diagnose ambiguous production failures from evidence and deliver practical, durable fixes.
- Meaningful coding experience, strong ownership, and ability to lead scoped or multi-team AI systems work.
- Bonus: deep Rust, systems, sandboxing, isolation, or low-level platform experience.
- Bonus: experience with coding agents, agent harnesses, tool-using LLM systems, model evaluations, or post-training feedback loops.
- Bonus: background in compilers, kernels, runtimes, inference optimization, GPU systems, benchmarking, profiling, or performance engineering.
- Bonus: experience building production infrastructure for many engineers or users under demanding reliability and security constraints.
- Bonus: open-source infrastructure or developer-platform experience with strong API and usability judgment.
Categories
About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. AI is an extremely powerful tool that must be created with safety and human needs at its core. OpenAI is dedicated to putting that alignment of interests first — ahead of profit. To achieve our mission, we must encompass and value the many different perspectives, voices, and experiences that form the full spectrum of humanity. Our investment in diversity, equity, and inclusion is ongoing, executed through a wide range of initiatives, and championed and supported by leadership. At OpenAI, we believe artificial intelligence has the potential to help people solve immense global challenges, and we want the upside of AI to be widely shared. Join us in shaping the future of technology.