Blink Health

Senior Site Reliability Engineer

Blink Health
Apply
3 months ago
Delhi, IndiaSenior

Responsibilities

  • Establish and evolve organization-wide SRE best practices, including reliability principles, error budgets, incident response, postmortems, and operational readiness.
  • Define and drive observability strategy through SLIs, SLOs, alerting quality, dashboards, and service health indicators.
  • Design and implement software-driven infrastructure solutions that automate manual processes and reduce operational toil.
  • Influence priorities and technical decisions across cloud infrastructure, reliability tooling, and platform architecture.
  • Lead large, ambiguous initiatives from concept through delivery while aligning engineering, security, and product stakeholders.
  • Improve platform resilience, scalability, performance, and compliance through software, infrastructure, and security expertise.
  • Identify systemic risks and reliability gaps and lead platform upgrades and architectural improvements.
  • Improve developer workflows, tooling, and operational maturity in partnership with engineering teams.
  • Provide technical mentorship, architecture guidance, and design and code reviews.
  • Document systems and processes and support knowledge sharing across teams.
  • Mature incident response, escalation practices, and post-incident learning.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.
  • 10+ years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles with demonstrated impact at scale.
  • Expert troubleshooting across application, kernel, and network layers, with deep Linux and operating system expertise.
  • Advanced networking knowledge including load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication.
  • Experience with Python, Go, Bash, and troubleshooting application stacks such as React.
  • Experience automating complex operational work and building internal tools in Python or Go.
  • Deep experience with AWS, GCP, or Azure cloud platforms and production-grade managed services.
  • Strong Kubernetes and container orchestration expertise, including EKS and Helm.
  • Experience designing observability systems covering metrics, logging, tracing, dashboards, and alerting.
  • Understanding of container technologies, security scanning, secrets management, dynamic configuration, microservices architectures, service meshes, and traffic management.
  • Experience designing and maintaining company-wide infrastructure-as-code codebases using Terraform, Pulumi, CloudFormation, or Ansible.
  • Ability to balance infrastructure cost, reliability, security, and long-term maintainability.

Benefits

  • Opportunity to work on healthcare technology products serving millions of patients.
  • Highly collaborative, cross-functional environment focused on learning and innovation.
  • Equal opportunity employer committed to diversity.

Categories

DevOpsSite Reliability
Contact me