3 months ago
Base Salary
$200k - $400k/yr
Responsibilities
- Own the reliability and scalability of production systems handling growing social data volumes and real-time AI workloads.
- Define and drive SLOs, SLIs, and error budgets.
- Build observability, alerting, and on-call practices.
- Lead incident response and blameless postmortems and implement systemic improvements.
- Improve performance, cost efficiency, and capacity planning across cloud infrastructure.
- Harden infrastructure-as-code, deployment, and CI/CD pipelines for resilience and repeatability.
- Partner with engineering teams to embed reliability into system design and raise operational standards.
Requirements
- At least 5 years of experience operating production systems as an SRE, infrastructure engineer, or platform engineer.
- Experience scaling databases, data infrastructure, or complex production platforms under significant load.
- Hands-on expertise with AWS or similar cloud infrastructure and infrastructure-as-code tooling.
- Strong programming skills for automation, tooling, and operational services.
- Comfort working in a fast-moving startup environment with high ownership and autonomy.
- Experience establishing or maturing an SRE practice at an early-stage or rapidly scaling company is a bonus.
- Familiarity with AWS, Pulumi, Postgres, ClickHouse, Turbopuffer, or Temporal is a bonus.
- Background in capacity planning, performance engineering, or cost optimization at scale is a bonus.
Benefits
- Health, vision, and dental benefits with a 401(k) match.
- Free lunch in Palo Alto.
- Competitive compensation and early equity.
- Career growth opportunities as the company scales.
- Exposure to AI tooling and the opportunity to shape AI-native marketing infrastructure.
Tech Stack
Categories
Site Reliability
