2 months ago
Base Salary
$170k - $300k/yr
Responsibilities
- Own the compute platform, including container orchestration, networking, autoscaling, service topology, and ingress.
- Own the data layer across Postgres, OpenSearch, and BigQuery, including performance, scaling, reliability, and cost effectiveness.
- Build cloud-spend attribution and optimization across compute, data, and storage.
- Define platform reliability practices, including SLOs, capacity planning, infrastructure observability, and incident response.
- Build infrastructure-as-code, service patterns, and self-service primitives for product teams.
Requirements
- Experience operating production Kubernetes at meaningful scale, including networking, autoscaling, and cluster reliability.
- Deep AWS expertise, including EKS, VPC, and IAM boundaries.
- Strong Terraform and infrastructure-as-code experience focused on platforms and reusable primitives.
- Experience completing a low-downtime database migration and operating a relational database such as Postgres under production load.
- Ability to reason from query plans to capacity plans and simplify complex systems.
- Comfort operating in a regulated, high-stakes environment; familiarity with PHI, HIPAA, or SOC 2 is a strong plus.
- Demonstrated ability to move quickly, find shortcuts, and deliver ambitious goals.
Benefits
- Fully covered medical, vision, and dental insurance.
- Memberships for One Medical, Talkspace, Teladoc, and Kindbody.
- Unlimited paid time off and 16 weeks of parental leave.
- 401(k) plan, FSA option, commuter benefits, and DashPass.
- Lunch at the office every day and dinner at the office after 6:30 p.m.
- Meaningful equity for full-time employees.
- Fully in-person work in New York.
