5 months ago
Responsibilities
- Design and build scalable, fault-tolerant infrastructure systems across multiple cloud regions.
- Own and evolve multi-cloud infrastructure on Azure and GCP, including Kubernetes orchestration, networking, and container management.
- Lead observability, incident response, and operational excellence initiatives.
- Architect distributed reliability systems involving load balancing, quota management, and failover mechanisms.
- Partner with Product Engineering and Security teams to support reliable and secure platform development.
- Drive reproducible and auditable infrastructure deployments using Terraform and Pulumi.
- Design a model proxy architecture for millions of daily inference requests and seamless model integration.
- Build Redis-backed distributed rate limiting and quota management systems.
- Architect multi-region deployments that satisfy data residency requirements.
- Develop observability infrastructure with SLA monitoring, burn rate alerts, and token attribution for cost tracking.
- Lead the evolution of CI/CD pipelines while maintaining production stability.
- Mentor engineers and raise technical standards through code reviews, design reviews, and technical leadership.
Requirements
- 8+ years of experience in infrastructure engineering or platform engineering in a production environment.
- A long track record of building and scaling complex, large-scale distributed systems.
- Deep proficiency with cloud infrastructure platforms, with Azure preferred and GCP or AWS experience accepted.
- Strong fluency with Terraform, Pulumi, or CloudFormation for infrastructure as code.
- Strong understanding of Kubernetes, container orchestration, networking, and cloud security at scale.
- Experience with Datadog, Sentry, PagerDuty, and Incident.io, as well as incident response practices.
- Strong programming skills in Python, Go, or similar languages.
- Experience with AI/ML infrastructure or high-throughput inference systems is preferred.
- Experience with distributed rate limiting, load balancing, quota management, multi-tenant platforms, security, and compliance is preferred.
- A track record of leading complex cross-functional projects and delivering measurable impact is preferred.
- Applicants must be authorized to work in India; visa sponsorship is not available.
Benefits
- The role is based in Bengaluru, India.
- Applicants must be authorized to work in India; visa sponsorship is not available.
- Harvey is an equal opportunity employer and provides reasonable accommodations for applicants with disabilities.
Tech Stack
Categories
DevOpsSite Reliability
About Harvey
Harvey is domain-specific AI for legal and professional services. Adopted by Fortune 500 companies like AT&T, Verizon, Cox, Koch, KKR, Bridgewater, more than 100,000 lawyers across 2,400+ customers in 70 countries and over 75% of AmLaw 100 law firms rely on Harvey to advance legal expertise faster across contract analysis, due diligence, compliance, and litigation. Backed by Sequoia Capital, OpenAI, GV, Kleiner Perkins, Coatue and EQT, Harvey is the trusted partner in modernizing the legal industry.
