
Senior Site Reliability Engineer
Clearwater Analytics3 months ago
Mumbai, IndiaSenior
Responsibilities
- Build Python-based internal tools and automation for monitoring, diagnosing, and supporting client deployments across AWS and Azure.
- Detect and remediate configuration and infrastructure drift and standardize client environments around shared golden paths.
- Build fleet-wide monitoring, alerting, and dashboards to improve deployment observability.
- Convert manual runbooks into automated checks, self-healing jobs, and one-click diagnostic and remediation tools.
- Extend Terraform-based client provisioning and deployment pipelines to make onboarding faster and more repeatable.
- Work with onboarding, support, and client success teams to identify and reduce operational toil.
Requirements
- 7-10 years of experience in software engineering, site reliability engineering, DevOps, or platform engineering.
- Strong Python programming skills, including production-quality code and tests.
- Hands-on experience with AWS or Azure cloud infrastructure, including networking, IAM/RBAC, storage, and compute.
- Working knowledge of infrastructure-as-code, ideally Terraform, and managing multiple environments through shared modules and per-environment configuration.
- Solid Linux fundamentals, including log analysis, process tracing, service debugging, and automation.
- Experience operating multi-tenant or fleet-style environments is preferred.
- Experience with observability stacks, including metrics, log aggregation, alerting, and dashboards, is preferred.
- Formal incident management experience, including on-call, postmortems, and blameless root-cause analysis, is preferred.
- Exposure to financial services, fintech, or other regulated environments is preferred.