6 months ago
Melbourne, Australia or London, United KingdomSenior
Responsibilities
- Participate in on-call rotations and incident response, including service restoration and incident communication.
- Identify reliability risks and recurring issues and drive improvements through alerting, automation, system changes, and process improvements.
- Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services.
- Strengthen dashboards, alerts, logs, and traces to improve detection and diagnosis.
- Automate repetitive tasks, simplify runbooks, and improve operational tooling.
- Improve deployments, rollback mechanisms, and operational readiness.
- Write and maintain runbooks, participate in blameless post-mortems, and improve incident response practices.
- Collaborate with product and feature teams on production readiness, service ownership, and reliability expectations.
Requirements
- 3–6+ years of experience in SRE, DevOps, platform, or operations-heavy engineering roles.
- Experience supporting production systems and participating in on-call rotations.
- Ability to debug live systems under pressure.
- Experience operating cloud infrastructure, with AWS preferred.
- Working knowledge of Kubernetes and containerized workloads.
- Infrastructure as Code experience with Terraform or a similar tool.
- Familiarity with monitoring and alerting tools such as Datadog and Prometheus.
- Scripting or automation experience with Python, Bash, or similar tools.
- Experience leading incidents or mentoring others during on-call is preferred.
- Experience in regulated or security-sensitive environments is preferred.
- Familiarity with production databases, queues, and caches is preferred.
- Interest in SLOs, error budgets, and capacity planning is preferred.
Benefits
- Equity from day one.
- Personal development budget.
- Ability to work from anywhere for one month.
- Dedicated wellness days and birthday off.
- Hybrid work environment with 3 days in the office.
Tech Stack
Categories
DevOpsSite Reliability
About Heidi
Heidi builds an AI care partner for clinicians that automates documentation and workflows, including an AI medical scribe, Evidence for point‑of‑care research, and Comms for patient coordination. It sells a freemium and enterprise SaaS platform to healthcare providers and health systems. Headquartered in Melbourne and privately held, Heidi reports supporting over 2.7 million patient interactions each week in 110 languages across 190 countries.
