
Principal Site Reliability Engineer
ServiceTitan6 days ago
Bengaluru, IndiaStaff+
Responsibilities
- Define the technical vision and long-term architecture for the reliability and infrastructure platform.
- Lead major platform initiatives involving Kubernetes infrastructure, observability, SLO programs, and incident management systems.
- Review architecture and infrastructure decisions with engineering and product leadership and establish organization-wide non-functional requirements.
- Identify systemic reliability risks and drive remediation across teams.
- Define and operationalize SLIs, SLOs, and error budgets across the engineering organization.
- Drive adoption of reliability and observability practices through documentation, design reviews, and partnership with product teams.
- Design and build AI-assisted operational systems for telemetry correlation, failure diagnosis, remediation, investigation, triage, and deployment safety.
- Partner with Infrastructure Engineering on progressive delivery, canary analysis, and AI-assisted deployment promotion decisions.
- Mentor Staff and Senior SREs through code reviews, architecture feedback, and pairing on complex problems.
- Participate in system design interviews and help define the Principal-level technical hiring bar.
Requirements
- 12+ years of relevant hands-on experience.
- Prior Staff or Principal-level experience with ownership of a technical domain.
- Expert-level Kubernetes knowledge, including internals, failure modes, capacity planning, and large-scale cluster management.
- Experience defining and operationalizing SLIs, SLOs, and error budgets across an engineering organization.
- Deep expertise with modern observability stacks such as OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
- Expert grounding in AWS, GCP, or Azure, including networking, security, and cost optimization at scale.
- Strong distributed-systems knowledge, including complex failure modes and graceful degradation.
- Experience leading major incident responses, conducting blameless postmortems, and implementing systemic fixes.
- Deep experience with CI/CD systems, progressive delivery, and pipeline reliability and safety.
- Production experience building or deploying AI-assisted operational capabilities such as automated investigation, anomaly correlation, remediation agents, or deployment safety systems.
- A track record of technical leadership on complex, cross-team infrastructure initiatives at product companies operating at scale.
- Ability to set technical direction, make architectural trade-offs, and build consensus across teams without formal authority.
Tech Stack
Categories
DevOpsSite Reliability
About ServiceTitan
ServiceTitan builds a cloud platform for home and commercial service contractors to run operations, including CRM, scheduling/dispatch, field mobile apps, estimates, invoicing, payments, and reporting. It also offers embedded fintech for payments and consumer financing, plus integrations for inventory and accounting. Founded in 2012 and headquartered in Glendale, California, it serves contractors across HVAC, plumbing, electrical, and roofing, and sells via subscription with add-on modules.