
Site Reliability Engineer
Clearwater Analytics3 months ago
Mumbai, IndiaMid Level / Senior
Responsibilities
- Build Python tools and automation to monitor, diagnose, and support client deployments across AWS and Azure.
- Detect and remediate configuration and infrastructure drift and migrate legacy deployments to standardized golden paths.
- Build monitoring, alerting, and dashboards to improve fleet-wide observability.
- Convert operational runbooks into automated checks, self-healing jobs, and one-click remediation tools.
- Extend Terraform-based client provisioning and deployment pipelines.
- Partner with onboarding, support, and client success teams to identify and reduce operational toil.
Requirements
- 3–5 years of experience in software engineering, site reliability engineering, DevOps, or platform engineering.
- Strong Python programming skills and experience writing production-quality code with tests.
- Hands-on experience with AWS or Azure, including networking, IAM/RBAC, storage, and compute.
- Working knowledge of infrastructure-as-code, ideally Terraform, and managing shared modules and per-environment configuration.
- Solid Linux fundamentals, including log analysis, process tracing, service debugging, and automation.
- Experience with multi-tenant or fleet-style environments is a plus.
- Observability stack experience involving metrics, log aggregation, alerting, and dashboards is a plus.
- Formal incident management experience, including on-call, postmortems, or blameless RCA practices, is a plus.
- Exposure to financial services, fintech, or other regulated environments is a plus.
Benefits
- Direct, visible impact on client onboarding and support efficiency.
- Broad exposure to cloud infrastructure, a large Python platform codebase, deployment pipelines, and operational workflows.
- Growth opportunities alongside experts in cloud infrastructure, quantitative finance, and large-scale SaaS operations.