about 2 hours ago
Base Salary
$182k - $242k/yr
Responsibilities
- Lead incident response efforts and coach junior team members.
- Document incidents and conduct root cause analysis.
- Develop and improve incident response playbooks.
- Build strategies for core service performance and supportability.
- Own system observability using tools like Prometheus and Grafana.
- Lead automation efforts for incident detection and recovery.
- Define KPIs and SLAs for incident management.
- Collaborate with engineers to improve platform reliability.
- Design and implement solutions for operational efficiency.
- Document hardware automation workflows and processes.
- Create CI/CD pipelines and ensure smooth server hardware lifecycle.
- Partner with Fleet Operations to design scalable tooling.
- Build dashboards and alerts for operational troubleshooting.
- Participate in on-call rotation and triage support issues.
Requirements
- 7+ years of experience in cloud operations or site reliability engineering.
- Understanding of cloud platforms like Kubernetes, AWS, and GCP.
- Familiarity with incident management practices and frameworks.
- Proficiency in Go programming language.
- Experience with Prometheus and Grafana.
- Prior experience deploying containerized applications using Kubernetes.
- Excellent documentation skills and attention to detail.
- Strong analytical and problem-solving abilities.
- Experience in on-call rotation supporting production services.
Benefits
- 100% paid medical, dental, and vision insurance for employees.
- Company-paid life insurance and short/long-term disability insurance.
- Flexible Spending Account and Health Savings Account.
- Tuition reimbursement and employee stock purchase program participation.
- Mental wellness benefits and family-forming support.
- Paid parental leave and flexible childcare support.
- 401(k) with a generous employer match.
- Flexible PTO and catered lunch each day.
- Casual work environment focused on innovative disruption.
Tech Stack
Categories
Data EngineeringDevOps
