about 3 hours ago
Base Salary
$161k - $242k/yr
Responsibilities
- Design, build, and operate the production infrastructure for Harvey’s products and AI workloads.
- Drive technical direction across compute infrastructure, networking, and Kubernetes.
- Lead cross-functional technical initiatives to enhance reliability, scalability, and operational efficiency.
- Partner with various teams to translate product requirements into resilient infrastructure solutions.
- Establish reusable patterns and tooling for safe production service operations.
- Build and operate global compute and network infrastructure ensuring high availability and performance.
- Develop capacity models and automate fleet lifecycle management.
- Improve the Kubernetes platform and drive infrastructure cost efficiency.
- Build secure infrastructure foundations and develop Infrastructure-as-Code frameworks.
- Participate in on-call rotation and lead incident response efforts.
Requirements
- 5+ years of experience in software, infrastructure, site reliability, or production engineering.
- Deep experience with large-scale cloud infrastructure on AWS, Azure, or Google Cloud Platform.
- Strong hands-on experience with Kubernetes in production environments.
- Experience with distributed systems focusing on reliability and scalability.
- Proficiency in infrastructure automation using tools like Terraform or Pulumi.
- Strong understanding of compute infrastructure, networking, and production operations.
- Experience designing observability systems for monitoring and incident response.
- Knowledge of infrastructure security best practices.
- Proven ability to drive complex technical initiatives and influence decisions.
- Excellent communication skills for explaining technical concepts.
