3 months ago
Base Salary
$200k - $270k/yr
Responsibilities
- Join the on-call rotation, review stability incidents, and convert incident patterns into a prioritized reliability backlog.
- Define and operate SLOs, error budgets, and SLO-based alerting for high-impact services.
- Lead incident command and postmortem processes while improving reliability culture and MTTR.
- Perform capacity planning, load testing, provider rate-limit auditing, and per-organization concurrency analysis.
- Operate Kubernetes production workloads, including pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, and graceful shutdown.
- Tune KEDA and custom-metrics autoscaling for wscaler and workerpool-cron-scaler.
- Build platform services such as capacity forecasters, auto-remediation systems, and on-call tooling in Go or TypeScript.
- Drive measurable improvements in p99 call completion and platform reliability.
Requirements
- Experience running incident command and postmortem processes on a real production on-call rotation.
- Experience operating SLOs and error budgets with Chronosphere, Prometheus, Grafana, or Datadog.
- Experience with capacity planning and load testing for production systems with real users.
- Fluency in Kubernetes production operations, including pod crash diagnosis, HPA/VPA tuning, PodDisruptionBudgets, and graceful shutdown.
- Knowledge of backpressure and autoscaling patterns, including KEDA and custom metrics scaling.
- Ability to build platform services in Go or TypeScript rather than only writing scripts.
- Experience with real-time or latency-sensitive products is preferred.
- Experience as an SRE or Production Engineer at a large-scale technology company or in a real-time product environment is a potential background.
Benefits
- Competitive salary and equity ownership.
- Medical, dental, and vision coverage.
- Quarterly company off-sites and team activities.
- Flexible time off.
- Catered meals, transportation, gym access, and a $10,000 annual learning and development budget.
Tech Stack
Categories
Site Reliability
