ServiceTitan

Senior Site Reliability Engineer

ServiceTitan
Apply
1 day ago
Bengaluru, IndiaSenior

Responsibilities

  • Participate in an on-call rotation and use runbooks and playbooks to diagnose and resolve production issues.
  • Design, build, and maintain observability dashboards and alerting based on Service Level Indicators and Service Level Objectives.
  • Operate and improve the Kubernetes-based compute platform.
  • Support reliable and scalable systems across Azure and AWS cloud networking and infrastructure.
  • Investigate production incidents, perform root-cause analysis, and complete remediation work.
  • Use AI-assisted engineering tools to automate investigations and ship fixes across infrastructure and application repositories.
  • Define scalability, availability, and performance requirements for new systems.
  • Review architecture and infrastructure decisions with product engineering teams.
  • Drive reliability and observability best practices across engineering teams.
  • Build automation that reduces repetitive operational work.
  • Maintain runbooks and documentation for on-call knowledge sharing.
  • Contribute to CI/CD pipelines and safe software delivery.

Requirements

  • Strong hands-on understanding of Kubernetes.
  • Practical experience defining and monitoring SLIs, SLOs, and error budgets on real systems.
  • Solid cloud engineering and networking fundamentals, including experience with AWS, GCP, or Azure and knowledge of subnetting and IP addressing.
  • Deep experience with at least one modern observability stack, such as OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
  • Strong understanding of a CI/CD system, with GitHub Actions preferred and TeamCity, Azure DevOps, or GitLab CI accepted.
  • Experience building web applications using the .NET stack, Python with Flask or FastAPI, or Java with Spring, deployed at scale.
  • Foundational experience supporting distributed systems and handling retry storms, timeouts, and cascading failures at scale.
  • Experience handling live-site incidents and troubleshooting under pressure, with the ability to improve TTx SLA.
  • Required experience using AI tools, including building SRE skills and deploying root-cause analysis agents in production systems.
  • 7+ years of relevant hands-on experience.
  • Ability to take accountability for reliability of business-critical, large-scale enterprise systems and make decisions with limited information.

Tech Stack

AWSAzureDatadogElasticsearchFastAPIFlaskGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaJavaKubernetes.NETPrometheusPythonTeamCity

Categories

Site Reliability
ServiceTitan

About ServiceTitan

1,001-5,000 employees

ServiceTitan builds a cloud platform for home and commercial service contractors to run operations, including CRM, scheduling/dispatch, field mobile apps, estimates, invoicing, payments, and reporting. It also offers embedded fintech for payments and consumer financing, plus integrations for inventory and accounting. Founded in 2012 and headquartered in Glendale, California, it serves contractors across HVAC, plumbing, electrical, and roofing, and sells via subscription with add-on modules.

Contact me