about 3 hours ago
Base Salary
$231k - $340k/yr
Responsibilities
- Design, build, and operate the production infrastructure for Harvey’s products and AI workloads.
- Drive technical direction across compute infrastructure, networking, Kubernetes, and production operations.
- Lead cross-functional technical initiatives to improve reliability, scalability, and operational efficiency.
- Partner with various teams to translate product requirements into resilient infrastructure solutions.
- Establish reusable patterns and tooling to help engineering teams operate production services safely.
- Build and operate global compute and network infrastructure ensuring high availability and performance.
- Develop capacity models and demand forecasts to scale infrastructure efficiently.
- Operate and improve Harvey’s Kubernetes platform, including monitoring and operational automation.
- Drive cost efficiency through capacity management and workload optimization.
- Build secure infrastructure foundations with compliance controls.
- Develop Infrastructure-as-Code frameworks using tools like Terraform and Pulumi.
- Improve observability and incident response across the infrastructure platform.
- Participate in on-call rotation and lead incident response efforts.
Requirements
- 10+ years of experience in software, infrastructure, site reliability, or production engineering.
- Deep experience with large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
- Strong hands-on experience operating Kubernetes in production environments.
- Experience with distributed systems that exhibit reliability and scalability.
- Proficiency in infrastructure automation and Infrastructure-as-Code using Terraform or Pulumi.
- Strong understanding of compute infrastructure, networking, and production operations.
- Experience designing observability systems for monitoring and incident response.
- Knowledge of infrastructure security best practices.
- Proven track record of driving complex technical initiatives.
- Excellent communication skills for explaining technical concepts to stakeholders.
- A systems-thinking mindset focused on building reliable and scalable infrastructure.
