15 days ago
Manchester, United KingdomStaff+
Responsibilities
- Define SLOs, SLIs, error budgets, instrumentation standards, and reliability investment priorities for Tier-1 customer-facing services.
- Design and operate incident command, severity management, war rooms, postmortems, on-call rotations, and incident follow-up processes.
- Select and establish the primary observability platform and create dashboards, alerts, runbooks, and operational KPIs.
- Contribute code and infrastructure as code, oversee CI/CD evolution, and administer secure and compliant Kubernetes platforms.
- Lead Sev-1 and Sev-2 incident response, chaos and failover exercises, and reliability validation.
- Architect and evaluate AIOps capabilities including anomaly detection, predictive alerting, automated root-cause analysis, and LLM-assisted triage.
- Coach an Engineer III SRE, design future SRE hires, and represent SRE in architecture, launch-readiness, and executive reviews.
- Partner with Security, Compliance, Architecture, Support, Platform Engineering, and managed service providers in a regulated healthcare environment.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent experience.
- 8+ years of software or platform engineering experience, including at least 4 years in SRE, DevOps, or platform reliability.
- At least 2 years of formal technical leadership, tech-lead, or staff-level experience with mentorship responsibilities.
- Proven experience building an SRE, DevOps, or platform reliability practice from zero or near-zero, including SLOs, incident command, and error budgets.
- Deep hands-on expertise with AWS, Azure, or GCP, including networking, IAM, and managed services.
- Strong CI/CD pipeline design and management experience and infrastructure-as-code experience with Terraform, Chef, Puppet, or similar tools.
- Proficiency in Python or another object-oriented programming language and experience administering and scaling Kubernetes clusters.
- Experience with observability platforms, real customer-impacting incident command, blameless postmortems, and follow-up discipline.
- Experience integrating AI/ML-based anomaly detection, alerting, or LLM-assisted triage, or strong judgment about AIOps in regulated environments.
- Ability to coach and mentor engineers, influence senior technical or operations leaders, communicate effectively, and operate under HIPAA and SOC 2 requirements.
- Preferred qualifications include a master’s degree, healthcare or regulated-industry experience, managed service provider integration, hybrid hardware-plus-cloud experience, AIOps platform experience, large language model APIs or agentic AI frameworks, stateful Kubernetes services, security scanning, intrusion detection, Kafka, RabbitMQ, chaos engineering, Databricks, Team Foundation Server, Octopus Deploy, and FinOps.
Benefits
- Corporate office, remote, or hybrid work arrangement is supported.
- Up to 10% travel is required.
- On-call participation is expected as part of the SRE rotation.
- The role offers a player-coach path toward Manager or Director of SRE or a Principal/Distinguished Engineer path.
Tech Stack
Apache KafkaAWSAzureChefCodefreshDatabricksDatadogDockerElasticsearchGitHub ActionsGoogle Cloud PlatformGrafanaHelmIstioJenkinsKibanaKubernetesOctopus DeployPrometheusPuppetPythonRabbitMQTeamCityTerraform
Categories
Site Reliability
About Omnicell
Omnicell is transforming pharmacy and nursing care through outcomes-centric solutions designed to optimize clinical and business outcomes across all settings of care. Our comprehensive portfolio of robotics and smart devices, intelligent software workflows, and data and analytics, all optimized by expert services are helping healthcare facilities worldwide to reduce costs, improve labor efficiency, establish new revenue streams, enhance supply chain control, support compliance, and move closer to the industry vision of the Autonomous Pharmacy. To learn more, visit omnicell.com.
