
Senior Site Reliability Engineer II
Blink Health15 hours ago
Pittsburgh, PA, USA +2 moreSenior
Responsibilities
- Define and drive observability strategy for IT system and process health, performance, reliability, alerting quality, dashboards, and service health indicators.
- Design and implement software-driven IT solutions that automate manual processes and reduce operational toil.
- Act as a technical leader across support, systems, networks, engineering, and partner teams.
- Own large, ambiguous initiatives from concept through delivery while aligning stakeholders.
- Identify systemic reliability risks, lead platform upgrades, and drive architectural improvements.
- Provide technical mentorship, architecture guidance, and design and code reviews.
- Improve documentation and knowledge sharing across systems and processes.
- Participate in and mature incident response, escalation practices, and post-incident learning.
Requirements
- Bachelor’s or Master’s degree in Computer Science or equivalent practical experience.
- At least 5 years of experience in site reliability engineering, infrastructure engineering, or platform engineering roles with demonstrated impact at scale.
- Expert troubleshooting across applications, kernels, operating systems, and networks, with deep Linux and command-line expertise.
- Advanced understanding of load balancing, proxies, DNS, TCP/IP, NAT, and service-to-service communication.
- Strong proficiency in at least one of Python, Go, or Bash, with familiarity troubleshooting application stacks such as React.
- Demonstrated ability to automate complex operational work and build internal tools in Python or Go.
- Experience with AWS or comparable cloud platforms including GCP or Azure and production-grade managed services.
- Expertise with Kubernetes and container orchestration, including EKS and Helm.
- Experience designing and implementing observability systems covering metrics, logging, tracing, dashboards, and alerting.
- Understanding of container technologies, security scanning, secrets management, dynamic configuration, microservices architectures, service meshes, and advanced traffic management.
- Experience maintaining company-wide infrastructure-as-code codebases using Terraform, Pulumi, CloudFormation, or Ansible.
- Ability to balance infrastructure cost, reliability, security, and long-term maintainability.
- Comfort working in an agile environment with disciplined testing and quality practices.
Tech Stack
Categories
DevOpsSite Reliability