
Lead Site Reliability Engineer, Incident Management
Qualys, Inc.2 hours ago
Pune, IndiaStaff+
Responsibilities
- Lead enterprise-wide critical incident response and serve as Incident Commander.
- Own end-to-end restoration strategies for complex production outages.
- Drive executive communications, customer-impact assessments, root cause analysis, and systemic reliability improvements.
- Design self-healing platforms and automation and improve observability, SLOs, SLIs, and error budgets.
- Lead capacity planning, operational readiness, architecture reviews, and cross-functional operational standards.
- Mentor Senior and Lead SREs and drive SRE strategy, reliability roadmaps, and reliability culture.
Requirements
- Bachelor's degree or equivalent.
- 8–12+ years of experience in SRE or production operations.
- Strong experience with Linux, Kubernetes, cloud, networking, and databases.
- Expertise in Python, Go, Bash, or Java.
- Experience leading major incidents and mentoring engineers.
- Excellent communication and leadership skills.
- Preferred qualifications include large-scale SaaS experience, multi-cloud expertise, Chaos Engineering experience, Cloud/Kubernetes certifications, and ITIL certification.
Benefits
- Full-time position with hybrid/remote work options.
- 24x7 production support and on-call leadership responsibilities.
- Cross-functional collaboration across engineering, infrastructure, product, security, support, and executive leadership.
Tech Stack
Categories
Site Reliability