
Senior Staff Site Reliability Engineer
ServiceTitan1 day ago
Bengaluru, IndiaStaff+
Responsibilities
- Participate in on-call rotations and diagnose and resolve production issues using runbooks and playbooks.
- Design, build, and maintain observability dashboards and alerting based on SLIs and SLOs.
- Operate and improve the Kubernetes-based compute platform.
- Support reliable, scalable systems across AWS and Azure cloud networking and infrastructure.
- Investigate production incidents, perform root-cause analysis, and complete remediation work.
- Use AI-assisted engineering tools and build production SRE agents to automate investigations and infrastructure fixes.
- Define scalability, availability, and performance requirements for new systems.
- Review architecture and infrastructure decisions with product engineering teams.
- Drive reliability and observability best practices across engineering teams.
- Build automation that reduces repetitive operational work.
- Maintain runbooks and documentation for shared on-call knowledge.
- Contribute to CI/CD pipelines and safe software delivery practices.
Requirements
- 10+ years of relevant hands-on experience.
- Strong hands-on understanding of Kubernetes.
- Practical experience defining and monitoring SLIs, SLOs, and error budgets.
- Cloud engineering and networking experience with AWS, GCP, or Azure, including subnetting and IP addressing.
- Deep experience with at least one modern observability stack such as OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
- Strong understanding of a CI/CD system, with GitHub Actions preferred; TeamCity, Azure DevOps, or GitLab CI are acceptable alternatives.
- Experience building web applications at scale using the .NET stack, Python with Flask or FastAPI, or Java with Spring.
- Experience supporting distributed systems and diagnosing failure modes such as retry storms, timeouts, and cascading failures.
- Experience handling live-site incidents and improving time-to-resolution service-level performance.
- Required experience using AI tools, including building production root-cause-analysis agents.
- Ability to guide decisions with limited information and own reliability across diverse problem areas.
Tech Stack
AWSAzureDatadogElasticsearchFastAPIFlaskGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaJavaKubernetes.NETPrometheusPythonTeamCity
Categories
Site Reliability
About ServiceTitan
ServiceTitan builds a cloud platform for home and commercial service contractors to run operations, including CRM, scheduling/dispatch, field mobile apps, estimates, invoicing, payments, and reporting. It also offers embedded fintech for payments and consumer financing, plus integrations for inventory and accounting. Founded in 2012 and headquartered in Glendale, California, it serves contractors across HVAC, plumbing, electrical, and roofing, and sells via subscription with add-on modules.