1 day ago
Hyderābād, IndiaSenior
Responsibilities
- Design, build, and maintain automation and tooling that reduces operational toil.
- Mature SRE practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination.
- Participate in detection, escalation, mitigation, stakeholder communication, post-incident reviews, and corrective-action follow-through.
- Participate in a 12-hour follow-the-sun on-call rotation within the Global SRE team.
- Use incident, support, alert, and SLO trends to drive proactive risk reduction and improve developer experience.
- Partner with service owners on production readiness, capacity planning, resilience testing, game days, disaster recovery, and workload modernization.
- Define service health expectations, document critical customer workflows, identify dependencies, and align reliability investments with business priorities.
- Build relationships with observability, incident response, and cloud infrastructure vendors.
- Apply AI-assisted and agentic workflows to alert triage, incident mitigation, postmortems, trend analysis, capacity planning, and SLO analysis with human oversight.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- At least 6 years of software engineering experience and 4 years of DevOps experience focused on CI/CD pipeline development.
- Production-facing SRE experience supporting a complex SaaS environment.
- Expertise with Kubernetes, Docker, orchestration tools, cloud platforms, infrastructure-as-code, observability tools, and scripting or programming languages.
- Experience establishing and maturing SRE principles, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination.
- Strong understanding of distributed systems, scalability, high availability, performance optimization, networking, and multi-cloud environments.
- Experience leading high-pressure incidents and communicating with technical teams, executives, customer-facing stakeholders, and vendors.
- Knowledge of event-driven autoscaling, advanced Kubernetes configurations, CI/CD tools, microservices, containerization, and DevOps practices.
- Strong problem-solving skills and the ability to work effectively in a fast-paced, collaborative environment.
Benefits
- Participate in a 12-hour follow-the-sun on-call rotation within a globally distributed SRE organization.
Tech Stack
Categories
DevOpsSite Reliability
About Seismic
Seismic builds an enablement and revenue execution platform that unifies governed sales content, buyer engagement signals, coaching, and AI-driven guidance for customer-facing teams. It sells its software by subscription to enterprises and mid-market companies, providing governed content, analytics, and coaching tools for go-to-market teams. Founded in 2010 and headquartered in San Diego, the privately held company serves about 2,500 organizations and 3.5 million users worldwide, with offices across North America, Europe, and Asia-Pacific.
