5 months ago
Base Salary
$230k - $342k/yr
Responsibilities
- Design, build, and operate networking systems supporting large-scale AI training and inference infrastructure.
- Improve performance, reliability, and scalability across host networking, datacenter fabrics, and WAN systems.
- Develop automation for provisioning, configuration management, validation, upgrades, and networking infrastructure lifecycle management.
- Build tooling and observability systems for network health, performance analysis, debugging, and automated remediation.
- Optimize network performance across RDMA, RoCE, InfiniBand, Ethernet, and high-performance GPU interconnects.
- Define and operationalize networking protocols, readiness criteria, and continuous validation systems.
- Partner with compute, storage, hardware, and infrastructure teams to scale networking with fleet growth.
- Contribute to topology design, capacity planning, failure-domain, and network-reliability architecture decisions.
- Diagnose complex distributed systems and networking issues across heterogeneous compute environments.
Requirements
- Experience building or operating large-scale networking or distributed systems infrastructure.
- Comfort working close to the hardware/software boundary.
- Experience with Linux networking, kernel systems, NICs, RDMA, or performance-sensitive infrastructure software.
- Experience with high-performance networking technologies such as InfiniBand, RoCE, DPDK, or large-scale Ethernet fabrics.
- Experience with datacenter networking, WAN systems, or host networking stacks.
- Ability to debug complex systems and performance bottlenecks across multiple layers of the stack.
- Production software development experience in C++, Python, or Go.
- Strong fundamentals in networking, operating systems, distributed systems, or infrastructure engineering.
About OpenAI
OpenAI builds and deploys large-scale AI models and tools—including ChatGPT, GPT-4–class models, DALL·E, and Whisper—sold via APIs and enterprise subscriptions to developers and businesses. It monetizes through usage-based API pricing and ChatGPT Plus/Team/Enterprise, and also reaches customers via Microsoft’s Azure OpenAI Service. Founded in 2015 and headquartered in San Francisco, it operates as a private partnership.
