2 months ago
Base Salary
$231k - $340k/yr
Responsibilities
- Design, build, and operate production infrastructure for Harvey’s products and AI workloads.
- Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
- Lead cross-functional initiatives improving reliability, scalability, security, operational efficiency, and infrastructure cost.
- Partner with Product Engineering, Security, AI Infrastructure, and Platform teams on resilient infrastructure solutions.
- Establish reusable patterns, tooling, and paved paths for safely shipping and operating production services.
- Build and operate global compute and network infrastructure with high availability, scalability, reliability, and performance.
- Develop capacity models, demand forecasts, and fleet lifecycle automation.
- Operate and improve the Kubernetes platform, including provisioning, upgrades, networking, monitoring, reliability, performance, and automation.
- Build secure infrastructure foundations covering identity and access management, network isolation, secrets management, auditing, and compliance controls.
- Develop Infrastructure-as-Code and automation frameworks using Terraform and Pulumi.
- Improve observability, monitoring, alerting, incident response, and operational readiness.
- Participate in on-call rotation, lead incident response when needed, document production learnings, and mentor engineers.
Requirements
- 10+ years of software, infrastructure, site reliability, or production engineering experience.
- Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
- Strong hands-on experience operating Kubernetes in production, including cluster lifecycle management, networking, and reliability.
- Experience building and operating distributed systems with strong reliability, scalability, and performance characteristics.
- Experience with infrastructure automation and Infrastructure-as-Code using Terraform or Pulumi.
- Strong understanding of compute infrastructure, networking, capacity planning, fleet management, and production operations.
- Experience designing and operating observability systems, including monitoring, logging, alerting, and incident response.
- Strong understanding of infrastructure security, including IAM, network security, secrets management, and compliance practices.
- Track record of driving complex cross-functional technical initiatives and influencing engineering decisions without formal authority.
- Clear communication skills, systems-thinking ability, and a focus on simple, reliable, scalable infrastructure platforms.
- Preferred experience supporting AI/ML or LLM infrastructure at scale, operating GPU fleets or high-performance computing infrastructure, managing multi-cloud or hybrid cloud environments, or building internal platforms and developer tooling.
Categories
DevOpsSite Reliability
About Harvey
Harvey builds domain-specific generative AI for legal and other professional services, delivered as an enterprise platform and APIs to automate contract analysis, due diligence, compliance, and litigation workflows. Founded in 2022 and headquartered in San Francisco, it sells to law firms and corporate legal departments on enterprise agreements; investors include Sequoia Capital and OpenAI. Customers include multiple Am Law 100 firms and Fortune 500 companies’ legal teams.
