
Senior Site Reliability Engineer
Blink Health3 months ago
Delhi, IndiaSenior
Responsibilities
- Establish and evolve organization-wide SRE best practices, including reliability principles, error budgets, incident response, postmortems, and operational readiness.
- Define and drive observability strategy through SLIs, SLOs, alerting quality, dashboards, and service health indicators.
- Design and implement software-driven infrastructure solutions that automate manual processes and reduce operational toil.
- Influence priorities and technical decisions across cloud infrastructure, reliability tooling, and platform architecture.
- Lead large, ambiguous initiatives from concept through delivery while aligning engineering, security, and product stakeholders.
- Improve platform resilience, scalability, performance, and compliance through software, infrastructure, and security expertise.
- Identify systemic risks and reliability gaps and lead platform upgrades and architectural improvements.
- Improve developer workflows, tooling, and operational maturity in partnership with engineering teams.
- Provide technical mentorship, architecture guidance, and design and code reviews.
- Document systems and processes and support knowledge sharing across teams.
- Mature incident response, escalation practices, and post-incident learning.
Requirements
- Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.
- 10+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles with demonstrated impact at scale.
- Expert troubleshooting across application, kernel, and network layers, with deep Linux and operating system expertise.
- Advanced networking knowledge including load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication.
- Experience with Python, Go, Bash, and troubleshooting application stacks such as React.
- Experience automating complex operational work and building internal tools in Python or Go.
- Deep experience with AWS, GCP, or Azure cloud platforms and production-grade managed services.
- Strong Kubernetes and container orchestration expertise, including EKS and Helm.
- Experience designing observability systems covering metrics, logging, tracing, dashboards, and alerting.
- Understanding of container technologies, security scanning, secrets management, dynamic configuration, microservices architectures, service meshes, and traffic management.
- Experience designing and maintaining company-wide infrastructure-as-code codebases using Terraform, Pulumi, CloudFormation, or Ansible.
- Ability to balance infrastructure cost, reliability, security, and long-term maintainability.
Benefits
- Opportunity to work on healthcare technology products serving millions of patients.
- Highly collaborative, cross-functional environment focused on learning and innovation.
- Equal opportunity employer committed to diversity.
Tech Stack
Categories
DevOpsSite Reliability