
Site Reliability Engineer (SRE)
TTEC Digital2 months ago
Hyderābād, IndiaSenior
Responsibilities
- Own SLOs and error budgets for each tenant and service.
- Lead incident response, incident command, blameless postmortems, and incident communications.
- Scale production systems and manage capacity modeling and load-testing partnerships.
- Build observability across metrics, logs, traces, latency percentiles, and alerting.
- Automate runbooks, remediation, and toil reduction using code.
- Operate and debug event-driven, real-time, WebSocket, and streaming systems.
- Establish chaos engineering and failure-injection practices to validate graceful degradation.
- Partner with DevOps on canary analysis, automatic rollback triggers, and error-budget-driven release gates.
- Monitor model latency, drift, and cost as production reliability signals.
- Participate in the on-call rotation and help establish 24/7 reliability coverage.
Requirements
- 8+ years operating production systems at scale, including ownership of SLOs, error budgets, and incident command.
- Strong Go or Python programming skills for reliability automation.
- Deep production experience with event-driven and real-time systems, streaming pipelines, WebSocket fleets, and failure modes involving state, ordering, locking, back-pressure, and cascading load.
- Strong monitoring and uptime mindset across metrics, logs, traces, and alerting.
- Good understanding of TCP/UDP, TLS, WebSocket, DNS, and load balancing; RTP/SIP experience is a strong plus.
- Experience operating GCP at scale; multi-cloud literacy and multi-tenancy isolation experience are strong pluses.
- Experience with capacity modeling, load testing, chaos engineering, canary analysis, and automatic rollback.
- Ability to communicate clearly and calmly during incidents with executives, customers, and technical teams.
- Strong production debugging skills using traces, metrics, and flame graphs.
Tech Stack
Categories
Site Reliability