5 hours ago
Base Salary
$230k - $385k/yr
Responsibilities
- Design and implement OS services, frameworks, and interfaces for inference execution, model loading, lifecycle management, and resource management.
- Partner with researchers to adapt and optimize models through quantization, runtime integration, and memory optimization for device constraints.
- Develop scheduling and resource policies that balance inference with other device activity while maintaining responsiveness and meeting latency, memory, battery, and thermal requirements.
- Develop and validate execution strategies for performance and power management across changing workloads and device conditions.
- Debug correctness, concurrency, performance, and reliability issues across models, inference runtimes, and OS components using tracing, profiling, and structured debugging.
- Build diagnostic tools, instrumentation, representative workloads, and automated tests to measure performance and energy improvements and catch regressions.
- Collaborate with research, hardware, firmware, platform, and product engineering teams to bring model capabilities into maintainable production systems.
Requirements
- Substantial hands-on experience designing, developing, and debugging operating-system components, system services, or performance-critical platform software.
- Proficiency in C++ systems development, including concurrent programming, memory ownership, and resource lifetime management.
- Hands-on experience integrating or optimizing inference runtimes or machine-learning workloads in resource-constrained environments.
- Strong understanding of scheduling, memory management, and competition for shared system resources.
- Experience diagnosing complex system behavior and delivering performance or power improvements supported by repeatable measurements.
- Ability to work across disciplines, translate research and product needs into system requirements, and clearly explain technical decisions and tradeoffs.
- Preferred qualifications include end-to-end on-device inference experience in shipped products, model adaptation through quantization or compression, inference optimization across CPUs, GPUs, or neural accelerators, and proficiency in Rust for systems programming.
- Preferred candidates take ownership across model, runtime, and operating-system boundaries and make tradeoffs among model quality, performance, power, reliability, security, and maintainability.
Categories
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
