Tandem Diabetes Care

Principal Site Reliability Engineer

Tandem Diabetes Care
Apply
1 day ago
Remote, United StatesStaff+

Base Salary

$165k - $185k/yr

Responsibilities

  • Lead day-to-day production support, incident command, escalation, stakeholder communication, postmortems, and corrective actions.
  • Own on-call rotation design, escalation paths, alert tuning, PagerDuty and New Relic tooling, and high-severity incident response.
  • Define and maintain SLIs and SLOs, observability, monitoring, alerting, runbooks, diagnostics, and automation to reduce MTTD and MTTR.
  • Lead Terraform-based infrastructure automation and eliminate repetitive operational toil.
  • Add automated rollback, change-risk checks, and progressive-delivery reliability guardrails to CI/CD pipelines.
  • Maintain supported, secure, and current cloud services, Kubernetes clusters, operating systems, runtimes, and infrastructure components.
  • Own backup, recovery, failover, recovery testing, runbooks, and RTO/RPO readiness for production platforms.
  • Maintain operational controls and audit evidence for access management, change management, logging, vulnerability remediation, patch management, security, and compliance.
  • Mentor and technically lead SRE, DevOps, consulting-partner, junior, and contract engineers through pairing, reviews, and incident debriefs.
  • Partner with software engineering, QA, architecture, Security, Quality, Compliance, and business stakeholders on reliability, resilience, regulatory compliance, and disaster recovery.
  • Support capacity planning, resilience testing, cloud cost optimization, and AI-enabled internal tooling and system integrations.

Requirements

  • Demonstrated experience leading production support and incident management, including incident command during high-severity events.
  • Strong grounding in SRE principles including SLIs, SLOs, blameless postmortems, toil reduction, and reliability engineering.
  • Experience owning on-call strategy, rotation design, alert tuning, and escalation.
  • Expertise with Terraform or comparable infrastructure as code at scale, including module design, state management, and policy-as-code guardrails.
  • Hands-on experience building reliable CI/CD pipelines with GitHub Actions, Octopus Deploy, or Azure DevOps.
  • Deep experience with AWS, Azure, or GCP and with Docker and Kubernetes.
  • Working knowledge of Prometheus, Grafana, Datadog, CloudWatch, ELK, or OpenSearch for SLO-based alerting.
  • Experience designing and testing disaster recovery, including backup and restore, failover, and RTO/RPO validation.
  • Working knowledge of cloud security, IAM, network segmentation, encryption, vulnerability management, patch management, and cloud cost optimization.
  • Proficiency in Python, Go, or Bash for automation and tooling.
  • 10+ years in Site Reliability Engineering, DevOps, or infrastructure engineering, including production support, infrastructure as code, CI/CD, disaster recovery, and security/compliance partnership.
  • At least 2 years mentoring or technically leading remote, offshore, or contracted engineers.
  • Bachelor of Science in Computer Science or equivalent education and applicable experience; technical school training and relevant certifications may also qualify, with production experience weighted more heavily than the degree.
  • Relevant AWS, Azure, or GCP Professional- or Architect-level certifications are preferred.
  • Experience in FDA- and ISO-regulated industries and with agile methodologies is preferred.

Benefits

  • Fully remote position open to candidates within the United States, with company-provided equipment and virtual training.
  • Medical, dental, and vision coverage available on the first day, plus health savings and flexible spending accounts.
  • 11 paid holidays and at least 20 days of paid time off with accrual beginning on day one.
  • 401(k) plan with company match and Employee Stock Purchase Plan.
  • Competitive bonus and benefits package in addition to base pay.

Tech Stack

AWSAzureBashDatadogDockerGitHub ActionsGoGoogle Cloud PlatformGrafanaKubernetesOctopus DeployPrometheusPythonTerraform

Categories

DevOpsSite Reliability
Tandem Diabetes Care

About Tandem Diabetes Care

1,001-5,000 employees

Tandem Diabetes Care designs and sells insulin delivery systems for people with diabetes, including the t:slim X2 and Tandem Mobi pumps with Control-IQ automated insulin delivery, plus the t:connect data app and infusion sets. The public company (NASDAQ) generates revenue from device sales and ongoing pump supplies. Founded in 2006 and headquartered in San Diego, it serves U.S. and international markets through healthcare providers and direct distribution.

Contact me