about 5 hours ago
Responsibilities
- Ensure reliability and performance of Plaud.ai’s AI products at scale.
- Design and operate highly available, scalable cloud-native systems for AI workloads.
- Own production reliability, incident response, and on-call practices.
- Build observability (metrics, logs, tracing) and reliability automation.
- Define and manage SLOs, SLIs, and error budgets with engineering teams.
- Drive postmortems and reliability improvements across the platform.
- Lead incident response and continuous reliability improvement.
- Partner with product and engineering teams on reliability design.
- Improve observability and operational maturity.
Requirements
- 5+ years in SRE, Infra, or Platform Engineering roles.
- Strong experience with cloud platforms (AWS/GCP/Azure/OCI).
- Hands-on with Kubernetes and distributed systems.
- Experience in on-call rotation and incident management.
- Proficient in at least one programming language (Go, Python, Java).
- Preferred: Experience supporting AI/ML or data-intensive platforms.
- Preferred: Experience of GPU cluster management.
- Preferred: Knowledge of SLO/SLA frameworks.
- Preferred: Experience in fast-growing or global products.
- Preferred: Exposure to multi-region systems.
- Preferred: Strong written and verbal communication.
Benefits
- Employee Stock Ownership Plan (ESOP) for long-term success.
- Top-tier medical, dental, and vision insurance with employer subsidy.
- 401(k) retirement plan with company matching for full-time employees.
- Unlimited PTO, 13 paid holidays, and 12 weeks of fully paid parental leave.
- Hybrid work model with a minimum of three in-office days per week.
- Access to high-quality office snacks, drinks, and equipment.
- Access to best-in-class AI tools for productivity.
- Choice of top-spec laptops and high-performance workstation setups.
- Annual company offsites and team events.
