Omnicell

Engineer III, Site Reliability

Omnicell
Apply
6 days ago

Responsibilities

  • Own reliability, scalability, instrumentation, alerting, dashboards, runbooks, and operational health for assigned cloud services.
  • Define and implement SLIs and SLOs with product and engineering teams and report reliability performance.
  • Identify operational toil and build automation to reduce repetitive manual work.
  • Participate in the SRE on-call rotation and progressively assume primary incident ownership.
  • Command Sev-2 and Sev-3 incidents, provide technical leadership during Sev-1 incidents, and lead blameless post-incident reviews.
  • Partner with IBM and HCL managed service providers on escalation from L1/L2 monitoring into SRE ownership.
  • Design, build, and operate CI/CD pipelines for cloud-native application delivery.
  • Automate infrastructure and platform services using Infrastructure as Code, preferably Terraform.
  • Improve observability through intelligent alerting, ML-based anomaly detection, and automated diagnostics.
  • Participate in architecture and launch readiness reviews and create reliable reference implementations and golden paths.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • At least 5 years of software or platform engineering experience, including at least 3 years in an SRE, DevOps, or reliability-focused role.
  • Hands-on experience with at least one major public cloud platform: AWS, Azure, or GCP.
  • Proficiency in Python or another object-oriented programming language for automation and tooling.
  • Production experience with Kubernetes, Docker, and Helm.
  • Experience implementing Infrastructure as Code with Terraform or similar frameworks.
  • Working knowledge of observability across metrics, logs, and tracing.
  • Real-world incident response experience, including on-call participation and post-incident write-ups.
  • Solid Linux system administration skills.
  • Preferred qualifications include regulated-environment experience, managed service provider familiarity, AIOps or ML-based anomaly detection exposure, GitOps knowledge, secure Kubernetes platform experience, and familiarity with chaos engineering, Kafka, RabbitMQ, or stateful Kubernetes services.

Benefits

  • Remote or hybrid work environment supported.
  • Up to 10% travel as needed.
  • Required participation in an SRE on-call rotation.
  • Hands-on mentorship from a Senior SRE in a player-coach environment.
  • Intentional growth path toward Senior Site Reliability Engineer, with potential lateral growth into platform engineering, security engineering, or product engineering.

Tech Stack

Apache KafkaAWSAzureCodefreshDockerGitHub ActionsGoogle Cloud PlatformHelmKubernetesLinuxOctopus DeployPythonRabbitMQTeamCityTerraform

Categories

Site Reliability
Omnicell

About Omnicell

1,001-5,000 employees

Omnicell is transforming pharmacy and nursing care through outcomes-centric solutions designed to optimize clinical and business outcomes across all settings of care. Our comprehensive portfolio of robotics and smart devices, intelligent software workflows, and data and analytics, all optimized by expert services are helping healthcare facilities worldwide to reduce costs, improve labor efficiency, establish new revenue streams, enhance supply chain control, support compliance, and move closer to the industry vision of the Autonomous Pharmacy. To learn more, visit omnicell.com.

Contact me