2 months ago
Base Salary
$231k - $340k/yr
Responsibilities
- Design, build, and operate production infrastructure for Harvey’s products and AI workloads.
- Drive technical direction across compute infrastructure, networking, Kubernetes, workflow orchestration, and production operations.
- Lead cross-functional initiatives improving reliability, scalability, security, operational efficiency, and infrastructure cost.
- Build and operate global compute and network infrastructure and improve utilization, performance, availability, and capacity planning.
- Operate and improve the Kubernetes platform, including provisioning, upgrades, networking, monitoring, reliability, performance, and automation.
- Build secure infrastructure foundations covering identity and access management, network isolation, secrets management, auditing, and compliance controls.
- Develop Infrastructure-as-Code and automation frameworks using Terraform and Pulumi.
- Improve observability, monitoring, alerting, incident response, and operational readiness, including participation in the on-call rotation.
- Establish reusable tooling and paved paths, conduct design reviews, document technical decisions, and mentor engineers.
Requirements
- 10+ years of software, infrastructure, site reliability, or production engineering experience is required.
- Deep experience building and operating large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform is required.
- Strong hands-on production Kubernetes experience, including cluster lifecycle management, networking, and reliability, is required.
- Experience with distributed systems, infrastructure automation, Infrastructure-as-Code using Terraform or Pulumi, observability, incident response, and infrastructure security is required.
- Candidates should understand compute infrastructure, networking, capacity planning, fleet management, IAM, network security, secrets management, and compliance practices.
- The role requires a track record of leading complex cross-functional technical initiatives and influencing engineering decisions without formal authority.
- Excellent communication, systems thinking, and an ability to build simple, reliable, and scalable infrastructure platforms are required.
- Experience supporting AI/ML or LLM infrastructure, operating GPU fleets or high-performance computing infrastructure, multi-cloud or hybrid cloud environments, or building internal platforms is preferred.
Categories
DevOpsSite Reliability
About Harvey
Harvey builds domain-specific generative AI for legal and other professional services, delivered as an enterprise platform and APIs to automate contract analysis, due diligence, compliance, and litigation workflows. Founded in 2022 and headquartered in San Francisco, it sells to law firms and corporate legal departments on enterprise agreements; investors include Sequoia Capital and OpenAI. Customers include multiple Am Law 100 firms and Fortune 500 companies’ legal teams.
