
Site Reliability Engineer
MyFitnessPal29 days ago
Remote, United StatesSenior
Base Salary
$120k - $165k/yr
Responsibilities
- Own and evolve SLI/SLO and error-budget frameworks to guide prioritization and product decisions.
- Lead incident response, postmortems, systemic remediation, and sustainable on-call practices.
- Build and maintain observability for metrics, logs, and traces using Datadog.
- Design and operate resilient infrastructure with Terraform, Kubernetes, and container workloads.
- Manage capacity planning, cloud-cost optimization, CI/CD pipelines, canary deployments, progressive rollouts, and fast rollback strategies.
- Integrate and tune SAST, DAST, SCA, dependency scanning, and policy-as-code controls in the software delivery pipeline.
- Implement admission-time policy enforcement using Kyverno, OPA/Rego, or Conftest.
- Drive vulnerability triage and remediation SLAs for pipeline and infrastructure findings.
- Build runbooks and automation, coach engineers, and promote reliability and operational best practices.
Requirements
- At least 5 years of experience in site reliability, platform, or infrastructure engineering with senior-level ownership of production systems.
- Strong programming skills in Go, Python, TypeScript, or a similar language, including experience building software or custom tooling.
- Hands-on experience with a major cloud platform, Kubernetes, and Infrastructure as Code; AWS and Terraform experience are preferred.
- Experience leading incident response and implementing SLO-driven reliability practices.
- Working fluency with observability tooling such as Datadog.
- Practical experience integrating security into CI/CD pipelines through SAST, DAST, SCA, dependency scanning, or policy-as-code.
- Strong understanding of cloud security fundamentals, including identity and IAM, least privilege, policy guardrails, and secrets management.
- Experience with admission-time policy-as-code frameworks, especially Kyverno; OPA/Rego and Conftest experience is also relevant.
- Experience in regulated or compliance-driven environments such as SOC 2, PCI DSS, or HIPAA is preferred.
- Chaos engineering, game-day, and high-traffic B2C/mobile backend experience are preferred.
- Strong judgment, communication, collaboration, and mentoring skills.
Benefits
- Comprehensive medical, dental, vision, healthcare, mental health, fertility, and family-support benefits.
- Paid maternity and paternity leave, responsible time off, and two volunteer days per calendar year.
- Annual performance bonus, 401(k) plan with employer match, and retirement savings program.
- Monthly wellness and technology allowances, mental health days, and access to MyFitnessPal Premium.
- Mentorship, virtual learning and development resources, training opportunities, recognition programs, and DEI initiatives.
- Teams meet in person as needed and the company gathers annually; the posting does not specify a fixed remote or office schedule.
Tech Stack
Categories
DevOpsSite Reliability