Blink Health

Senior Site Reliability Engineer II

Blink Health
Apply
15 hours ago
Pittsburgh, PA, USA +2 moreSenior

Responsibilities

  • Define and drive observability strategy for IT system and process health, performance, reliability, alerting quality, dashboards, and service health indicators.
  • Design and implement software-driven IT solutions that automate manual processes and reduce operational toil.
  • Act as a technical leader across support, systems, networks, engineering, and partner teams.
  • Own large, ambiguous initiatives from concept through delivery while aligning stakeholders.
  • Identify systemic reliability risks, lead platform upgrades, and drive architectural improvements.
  • Provide technical mentorship, architecture guidance, and design and code reviews.
  • Improve documentation and knowledge sharing across systems and processes.
  • Participate in and mature incident response, escalation practices, and post-incident learning.

Requirements

  • Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.
  • At least 5 years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles with demonstrated impact at scale.
  • Expert troubleshooting across applications, kernels, operating systems, and networks, with deep Linux and command-line expertise.
  • Advanced understanding of load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication.
  • Strong proficiency in at least one of Python, Go, or Bash, with familiarity troubleshooting application stacks such as React.
  • Demonstrated ability to automate complex operational work and build internal tools in Python or Go.
  • Experience with AWS or comparable cloud platforms including GCP or Azure and production-grade managed services.
  • Expertise with Kubernetes and container orchestration, including EKS and Helm.
  • Experience designing and implementing observability systems covering metrics, logging, tracing, dashboards, and alerting.
  • Understanding of container technologies, security scanning, secrets management, dynamic configuration, microservices architectures, service meshes, and advanced traffic management.
  • Experience maintaining company-wide infrastructure-as-code codebases using Terraform, Pulumi, CloudFormation, or Ansible.
  • Ability to balance infrastructure cost, reliability, security, and long-term maintainability.
  • Comfort working in an agile environment with disciplined testing and quality practices.

Categories

DevOpsSite Reliability
Contact me