ServiceTitan

Staff Site Reliability Engineer

ServiceTitan
Apply
8 days ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Participate in on-call rotations and diagnose and resolve production issues using runbooks and playbooks.
  • Design, build, and maintain observability dashboards and alerting based on SLIs and SLOs.
  • Operate and improve the Kubernetes-based compute platform.
  • Support reliable and scalable systems across AWS and Azure cloud networking and infrastructure.
  • Investigate production incidents, perform root-cause analysis, and complete remediation work.
  • Use AI-assisted engineering tools to automate investigations and ship fixes across infrastructure and application repositories.
  • Define scalability, availability, and performance requirements for new systems.
  • Review product engineering architecture and infrastructure decisions before release.
  • Drive adoption of reliability and observability best practices across engineering teams.
  • Build automation that reduces repetitive operational work.
  • Maintain runbooks and documentation for shared on-call knowledge.
  • Contribute to CI/CD pipelines and safe, efficient software delivery.

Requirements

  • Strong hands-on understanding of Kubernetes as a system.
  • Practical experience defining and monitoring SLIs, SLOs, and error budgets on real systems.
  • Solid cloud engineering and networking experience with AWS, GCP, or Azure, including subnetting and IP addressing.
  • Deep experience with at least one modern observability stack, such as OpenTelemetry, Prometheus, Grafana, Datadog, or Elasticsearch.
  • Strong understanding of a CI/CD system, with GitHub Actions preferred or equivalent experience with TeamCity, Azure DevOps, or GitLab CI.
  • Experience building web applications at scale using the .NET stack, Python with Flask or FastAPI, or Java with Spring.
  • Experience supporting distributed systems and handling failure modes such as retry storms, timeouts, and cascading failures.
  • Live-site incident handling and troubleshooting experience, including improving TTx SLA.
  • Required experience using AI tools for SRE skills, root-cause analysis agents, or similar production systems.
  • At least 9 years of relevant hands-on experience.

Tech Stack

AWSAzureDatadogElasticsearchFastAPIFlaskGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaJavaKubernetes.NETPrometheusPythonTeamCity

Categories

Site Reliability
ServiceTitan

About ServiceTitan

1,001-5,000 employees

We’re building the operating system for the trades, a critical industry that’s been underserved by technology for far too long. Founded by the sons of hard working tradespeople and backed by top investors, our platform delivers a seamlessly integrated experience that enables thousands of business owners to accelerate growth, drive operational efficiencies and deliver a superior customer experience. We currently serve over ten trades industries, and we’re just getting started. Joining our team means that you’ll have the opportunity to make an outsized impact on the trades ecosystem and world at large. Are you built for the challenge?

Contact me