
Lead Site Reliability Engineer - India
Juniper Square5 hours ago
Delhi, IndiaStaff+
Responsibilities
- Own the technical direction and architecture of team-owned infrastructure systems, balancing reliability, scalability, cost, and technical debt.
- Establish and maintain SLOs, monitor SLAs, improve observability, and lead incident response and root-cause analysis for complex multi-service issues.
- Define standards for logging, monitoring, operationalization, security, and production instrumentation.
- Lead the on-call rotation and implement preventative measures to protect customer experience and system reliability.
- Act as the directly responsible individual for medium-to-large SRE projects spanning multiple months and teams.
- Partner with Engineering Managers and Product Managers to scope roadmaps, plan delivery, negotiate scope, and resolve project risks.
- Mentor junior and mid-level engineers and drive technical alignment across engineering, product, and other stakeholders.
- Apply AI tools and operational AI systems to improve reliability, reduce MTTR, automate workflows, and enhance observability and incident response.
Requirements
- 7–10 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud environment.
- 5+ years of hands-on experience with AWS services across compute, networking, storage, and security.
- 5+ years managing Linux-oriented production environments at scale.
- 5+ years using Terraform, CDK, CloudFormation, and/or GitOps practices for Infrastructure-as-Code.
- 3+ years operating and troubleshooting production Kubernetes environments.
- 3+ years applying AWS Well-Architected Framework principles across reliability, security, performance, and cost.
- 3+ years working with cloud security practices including IAM, secrets management, network security, and compliance.
- 3+ years working with PostgreSQL in production, including performance tuning, replication, backup, and recovery.
- Strong programming skills and experience writing automation scripts or tooling in Python, Go, or similar languages.
- Deep knowledge of observability, distributed systems, data retention, backup and recovery, release management, and deployment automation.
- Familiarity with service mesh, API gateway patterns, and microservices architectures.
- Hands-on experience integrating LLMs or AI systems into production, including reliability, latency, observability, and failure handling.
- Knowledge of agent-based automation, LLMs, embeddings, RAG, model degradation, fallback strategies, and cost anomalies.
- Demonstrated ability to lead technical projects, make final technical decisions, communicate clearly, and mentor engineers.
Benefits
- Remote work or full-time, part-time, or occasional office work options are available, with offices in Mumbai and Bangalore for India-based employees.
- Employees collaborate through digital-first operations across India and other company locations.
- The role offers significant growth potential to help shape the SRE practice and prepare the platform for scale.
- The company provides a collaborative, engineering-driven culture focused on quality, curiosity, and ownership.
- Competitive compensation and benefits package.
Categories
Site Reliability
About Juniper Square
Juniper Square builds software and provides fund administration for private markets general partners, including real estate, private equity, and venture firms. Its SaaS platform supports fundraising, investor onboarding, compliance, treasury, reporting, and LP communications, with optional administration services and applied-AI tools integrated into clients’ systems. Founded in 2014 and headquartered in San Francisco, it is privately held and used by more than 2,000 GPs worldwide.