Sanity

Senior Site Reliability Engineer

Sanity
Apply
2 months ago
Remote, United StatesSenior

Responsibilities

  • Design, build, and operate shared platform foundations including GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability.
  • Diagnose and troubleshoot complex distributed systems operating at high request volume.
  • Analyze platform behavior and ensure effective observability.
  • Modernize edge, caching, and gateway layers onto Fastly and strengthen platform observability.
  • Improve dashboards, alert severity, paging standards, on-call readiness, and incident response.
  • Build golden paths, production-readiness checks, safe rollouts, and automation for safer deployments.
  • Mentor engineers and raise technical standards through code review, design review, and pairing.
  • Participate in the on-call rotation and support the rollout of developer on-call practices.

Requirements

  • 5+ years of experience participating in an SRE on-call rotation.
  • Experience with SRE/DevOps tools, processes, and culture.
  • Experience managing scalable, highly available, cloud-based applications with customer-facing uptime expectations.
  • Experience using Kubernetes to orchestrate, scale, and manage containerized cloud applications.
  • Experience building CI/CD pipelines and working with an observability stack such as Prometheus.
  • Experience across CDNs, edge systems, gateways, and caching layers, or willingness to develop deep expertise there.
  • Analytical infrastructure design, diagnosis, and optimization skills.
  • Comfort handling incidents and outages with thoughtful communication under pressure.
  • Applicants must be based in the United States with reasonable overlap with European engineering hours.

Benefits

  • Comprehensive health plans and perks.
  • Competitive stock options program and location-based salary.
  • Flexible, trust-based work environment supporting long-term professional and personal growth.
  • Healthy work-life balance accommodating individual and family needs.
  • Based in the United States with reasonable overlap with European engineering hours.

Tech Stack

Categories

DevOpsSite Reliability
Sanity

About Sanity

201-500 employees

Sanity is the intelligent content backend for companies building AI content operations at scale. Structured content that feeds models, powers agents, and runs workflows, not just websites. Used by Puma, Figma, Braze, Anthropic, and thousands of teams shipping content at scale. All-code Studio. Free to start.