about 3 hours ago
Base Salary
$236k - $290k/yr
Responsibilities
- Lead the design and implementation of Harvey's Model Infrastructure platform.
- Build systems to ensure high availability, low latency, and operational excellence for AI inference.
- Design and improve the Unified Model Controller (UMC) and Model Selector platform.
- Develop systems for model provisioning, capacity management, failover, and traffic engineering.
- Integrate new model providers and maintain provider APIs and SDKs.
- Improve observability through health dashboards, alerting, and analytics.
- Partner with Product Engineering to support model launches and monitoring.
- Drive infrastructure efficiency through capacity planning and cost visibility.
- Collaborate with AI Research for future model evaluation and deployment.
- Lead cross-functional technical initiatives and mentor engineers.
Requirements
- 7+ years of software engineering experience building large-scale distributed systems.
- Experience designing and operating highly available production services.
- Strong programming skills in Go, Java, Python, Rust, or C++.
- Deep understanding of distributed systems, cloud infrastructure, and observability.
- Experience leading technical projects across multiple engineering teams.
- Ability to balance long-term architecture with pragmatic execution.
- Strong communication and collaboration skills.
- Passion for building foundational platforms that enable other engineering teams.
