3 months ago
Responsibilities
- Design and maintain the core cloud infrastructure, deployment pipelines, and observability systems powering the real-time voice AI platform.
- Build internal tooling and abstractions that improve engineering velocity.
- Own reliability and performance at scale by defining SLOs, instrumenting monitoring, and leading incident response.
- Partner with backend engineers and researchers to translate AI capabilities into production-ready systems.
- Improve how the team builds, tests, and ships software across the development lifecycle.
Requirements
- Strong experience with cloud infrastructure using GCP, AWS, or Azure.
- Experience with containerized deployments using Docker and Kubernetes.
- Proficiency with infrastructure-as-code tooling such as Terraform or an equivalent.
- A track record of improving reliability and developer experience in high-throughput, latency-sensitive systems.
- Experience with real-time or streaming systems such as WebSockets or WebRTC is a plus.
- Familiarity with ML/AI infrastructure, including model serving, GPU workloads, or inference optimization, is a plus.
- Strong proficiency in TypeScript and Python is a plus.
- Being a former founder is a plus.
Benefits
- Free breakfast, lunch, and dinner provided in the office.
- Comprehensive health, dental, and vision coverage.
- Regular off-sites and team celebrations.
- Salary and equity are offered, with no specific compensation amount stated.
Categories
DevOpsSite Reliability
