5 months ago
Base Salary
$230k - $385k/yr
Responsibilities
- Design and build the core agent harness and execution loop for interpreting model outputs, using tools, executing code, and completing long-horizon tasks safely.
- Build sandboxing, isolation, orchestration, state, and workflow infrastructure for agents in real development environments.
- Develop evaluation, experimentation, and debugging systems that separate harness, model, inference/runtime, and product failures.
- Run ablations across prompts, model-facing interfaces, context construction, tool-use strategies, and harness behavior to improve solve rate, reliability, latency, and cost.
- Improve observability, profiling, and diagnostics across backend systems, inference, GPUs, and fleet capacity.
- Collaborate with research to make the harness trainable, measurable, and useful for improving frontier agentic models.
- Build shared primitives that make Codex faster, safer, more reliable, and easier for teams and open-source users to extend.
Requirements
- Production experience building or operating systems in distributed systems, infrastructure, developer tooling, sandboxing, virtualization, cloud platforms, or ML systems.
- Hands-on experience with LLM applications, coding agents, evaluation systems, model deployment, inference, compiler or runtime performance, or developer platforms.
- Ability to work across Rust systems code, Python configuration layers, APIs, agent orchestration, evaluations, logs and traces, inference behavior, runtime constraints, and user outcomes.
- Strong focus on reliability, safety, performance, debuggability, and clean abstractions.
- Ability to diagnose ambiguous production failures from evidence and deliver practical, durable fixes.
- Meaningful coding experience, strong ownership, and ability to lead scoped or multi-team AI systems work.
- Bonus: deep Rust, systems, sandboxing, isolation, or low-level platform experience.
- Bonus: experience with coding agents, agent harnesses, tool-using LLM systems, model evaluations, or post-training feedback loops.
- Bonus: background in compilers, kernels, runtimes, inference optimization, GPU systems, benchmarking, profiling, or performance engineering.
- Bonus: experience building production infrastructure for many engineers or users under demanding reliability and security constraints.
- Bonus: open-source infrastructure or developer-platform experience with strong API and usability judgment.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
