3 months ago
Tel Aviv-Yafo, IsraelSenior
Responsibilities
- Design, build, and maintain scalable, fault-tolerant systems.
- Define and enforce reliability processes, SLOs, SLIs, and SLAs.
- Lead complex incident responses, on-call rotations, and postmortems.
- Address observability, testing, production stability, and development productivity challenges.
- Drive reliability improvements through data-driven decisions.
- Build automation, tooling, and self-service capabilities.
- Collaborate with engineering, product, and support teams to embed reliability practices.
- Mentor engineers and promote operational excellence across the organization.
Requirements
- 7+ years of experience in SRE, DevOps, or Production Engineering roles, ideally in SaaS environments.
- Deep understanding of distributed systems, failure modes, resiliency patterns, observability, and large-scale production services running on Kubernetes.
- Hands-on experience building and owning monitoring tools.
- Experience with CI/CD tools.
- Proficiency with infrastructure-as-code tools.
- Solid experience with cloud platforms, preferably AWS.
- Experience with Java is an advantage.
Benefits
- Flexible hybrid work model combining working from home, on the go, or at the office.
Tech Stack
Categories
DevOpsSite Reliability
About Gong.io
Gong builds a revenue intelligence platform that captures and analyzes customer interactions (calls, emails, meetings) to deliver insights, coaching, and forecasting for sales and customer-success teams. It sells its software by subscription to businesses and integrates with common CRM and collaboration tools; more than 5,000 companies use it. Founded in 2015 and headquartered in San Francisco, Gong is a privately held company.
