
Senior Site Reliability Engineer
Teladoc HealthResponsibilities
- Define, implement, and improve SLIs, SLOs, and error budgets for critical applications and platform services.
- Improve service reliability, fault tolerance, scalability, operational readiness, and resilience against Azure, network, dependency, and deployment failures.
- Build observability across applications, infrastructure, networks, and cloud services using dashboards, alerts, logs, traces, and metrics.
- Analyze performance, capacity, saturation, bottlenecks, backup, disaster recovery, failover, and business continuity risks.
- Support and improve production workloads on Microsoft Azure and enforce operational standards for monitoring, backup, recovery, identity, security, and cost awareness.
- Lead incident response, blameless post-incident reviews, root-cause analysis, corrective actions, prevention plans, and runbook improvements.
- Support operational security and compliance controls for identity, access, encryption, secrets management, vulnerability remediation, logging, and auditability.
- Anchor and help establish the SRE practice, mentor the SRE team, and serve as the technical authority for observability and reliability.
Requirements
- At least 7 years of site reliability experience with hands-on ownership of mission-critical services through applicable work experience, training, military experience, or education.
- Deep Microsoft Azure experience, including Azure Monitor, Application Insights, AKS, and cloud-native operations across hybrid infrastructure.
- Proven experience designing and implementing production SLO programs with SLIs, SLOs, and error-budget policies.
- Hands-on experience with enterprise observability platforms such as Datadog, Dynatrace, Elastic, Grafana, Prometheus, or LogicMonitor.
- Hands-on experience configuring Datadog for monitoring, observability, and alerting.
- Preferred experience establishing or anchoring an SRE practice across multiple engineering teams and leading incident response and blameless postmortems.
- Preferred healthcare IT experience and familiarity with HIPAA, HITRUST, or equivalent compliance frameworks.
- Preferred AWS and multi-cloud reliability experience, chaos engineering, resiliency testing, infrastructure as code, scripting, programming, security controls, vulnerability management, compliance audits, and cloud governance.
- Preferred proficiency with Python, PowerShell, or Go and tools such as Gremlin, Chaos Mesh, Azure Chaos Studio, Terraform, Bicep, and Ansible.
- Recognized certifications such as Azure Solutions Architect, Google SRE certificate, or CKA are preferred.
Benefits
- Inclusive benefits programs centered on employees and their families, with tailored programs for individual needs.
- Opportunities for career growth, leadership, and meaningful work supporting access to virtual care.
- Innovative, inclusive workplace culture with equal employment opportunity and anti-discrimination protections.
Categories
About Teladoc Health
Teladoc Health is the global leader in virtual care, delivering and orchestrating care across patients, care providers, platforms and partners – transforming virtual care into a catalyst for how better health happens. Our comprehensive care offerings span primary care, mental health, chronic condition management and more, all integrated on a single platform. We partner with employers, health plans, hospitals, health systems and other providers to fuel clinical excellence and apply the power of technology to help people live their healthiest lives. Follow us for industry insights, virtual care trends, client success stories and expert perspectives shaping the future of healthcare. For more information, please visit www.teladochealth.com or follow @TeladocHealth on X.