7 months ago
Base Salary
$200k - $250k/yr
Responsibilities
- Design and build scalable, fault-tolerant infrastructure systems across multiple cloud regions.
- Own and evolve multi-cloud infrastructure on Azure and GCP, including Kubernetes orchestration, networking, and container management.
- Lead observability, incident response, and operational excellence initiatives.
- Architect and optimize distributed systems for reliability, including load balancing, quota management, and failover mechanisms.
- Partner with Product Engineering and Security teams to improve platform capabilities.
- Drive infrastructure-as-code practices with Terraform and Pulumi for reproducible, auditable deployments.
- Design a model proxy architecture routing millions of daily inference requests while maintaining model API compatibility.
- Build distributed rate-limiting and quota-management systems backed by Redis.
- Architect multi-region deployments that satisfy data residency requirements.
- Develop observability infrastructure with SLA monitoring, burn-rate alerts, and token attribution for cost tracking.
- Lead improvements to CI/CD pipelines while maintaining production stability.
- Mentor junior and intern engineers and provide technical leadership through code and design reviews.
Requirements
- At least 4 years of production experience in infrastructure engineering or platform engineering.
- Long track record building and scaling complex, large-scale distributed systems.
- Deep proficiency with cloud infrastructure platforms, with Azure preferred and GCP or AWS experience transferable.
- Strong fluency with infrastructure-as-code tools such as Terraform, Pulumi, or CloudFormation.
- Solid understanding of Kubernetes, container orchestration, networking, and cloud security at scale.
- Experience with observability tools such as Datadog and Sentry and incident response tools such as PagerDuty and Incident.io.
- Strong programming skills in Python, Go, or similar languages.
- Experience with AI/ML infrastructure or high-throughput inference systems is preferred.
- Experience with distributed rate limiting, load balancing, or quota management systems is preferred.
- Experience operating multi-tenant platforms with strict security and compliance requirements is preferred.
- Track record of leading complex cross-functional projects and delivering measurable impact is preferred.
- Strong problem-solving skills and commitment to operational excellence.
Benefits
- In-person work model in San Francisco, California
- Relocation assistance for new employees
Tech Stack
Categories
About Harvey
Harvey is domain-specific AI for legal and professional services. Adopted by Fortune 500 companies like AT&T, Verizon, Cox, Koch, KKR, Bridgewater, more than 100,000 lawyers across 2,400+ customers in 70 countries and over 75% of AmLaw 100 law firms rely on Harvey to advance legal expertise faster across contract analysis, due diligence, compliance, and litigation. Backed by Sequoia Capital, OpenAI, GV, Kleiner Perkins, Coatue and EQT, Harvey is the trusted partner in modernizing the legal industry.
