about 5 hours ago
Responsibilities
- Own the reliability architecture of the platform across regions and cloud providers.
- Collaborate with platform teams to provide support on operability and best practices.
- Set operational standards for on-call quality and incident response.
- Mentor and develop the SRE team technically.
- Participate in a 24/7 on-call rotation to resolve infrastructure issues.
Requirements
- 10+ years of experience in software and operating distributed systems.
- Deep expertise in Kubernetes, including multi-cluster platform design.
- Proficiency in Python, Go, or similar programming languages.
- Understanding of workload isolation at the systems level.
- Familiarity with AWS, GCP, or Azure infrastructure primitives.
- Strong preference for automation over manual processes.
Benefits
- Supportive and enriching culture for personal and professional growth.
- Employee affinity groups and fertility assistance.
- Generous parental leave policy.
- Commitment to providing accommodations for individuals with disabilities.