
Lead Site Reliability Engineer
Kontakt.io2 months ago
Base Salary
$200k - $250k/yr
Responsibilities
- Ensure 99.99% uptime and meet strict SLAs for healthcare customers.
- Design self-healing and fault-tolerant systems for a real-time cloud platform.
- Define SLIs, SLOs, and SLAs and improve monitoring and incident resolution.
- Architect and manage scalable AWS infrastructure for large-scale real-time data processing.
- Optimize Kubernetes and Docker environments for multi-region deployments.
- Lead infrastructure-as-code adoption with Terraform.
- Build monitoring, alerting, and logging systems using Prometheus, Grafana, OpenTelemetry, and Datadog.
- Lead incident response, on-call operations, postmortems, disaster recovery, and business continuity planning.
- Automate deployment, scaling, and failover to reduce manual intervention.
- Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
- Lead and mentor the SRE team and collaborate across Product, Engineering, Infrastructure, Security, and Compliance.
Requirements
- 10+ years of experience in Site Reliability Engineering or cloud infrastructure.
- Proven success scaling high-traffic, mission-critical platforms in SaaS, IoT, or healthcare.
- Deep expertise with AWS, Kubernetes, and distributed systems.
- Strong monitoring, logging, and observability experience with Prometheus, OpenTelemetry, or similar tools.
- Hands-on experience with incident management, postmortems, and resilient systems.
- Deep knowledge of infrastructure-as-code automation, Terraform, and GitOps.
- Ability to drive technical strategy and grow and mentor a high-performance SRE team.
- Understanding of network security, access management, HIPAA, and SOC 2.
- Bonus: healthcare IT experience including EHR data, FHIR, and HL7 interoperability.
- Bonus: experience with real-time distributed systems, event-driven architectures, large-scale data pipelines, on-call rotations, or major incident management.
Benefits
- Hybrid schedule requiring at least 3 days per week from the New York City office.
- Equity in a high-growth company backed by leading investors.
- Full health, dental, and vision coverage, 401k, paid time off, paid parental leave, and work tools.
- Autonomy to solve meaningful problems with quickly shipping work.
Tech Stack
Categories
DevOpsSite Reliability