Responsibilities
- Participate in on-call rotations and incident response, including service restoration and incident communication.
- Identify reliability risks and recurring issues and drive improvements through alerting, automation, system changes, and process improvements.
- Operate and improve Kubernetes clusters, cloud infrastructure, and core platform services.
- Strengthen dashboards, alerts, logs, and traces to improve detection and diagnosis.
- Automate repetitive tasks, simplify runbooks, and improve operational tooling.
- Improve deployments, rollback mechanisms, and operational readiness.
- Write and maintain runbooks, participate in blameless post-mortems, and improve incident response practices.
- Collaborate with product and feature teams on production readiness, service ownership, and reliability expectations.
Requirements
- 3–6+ years of experience in SRE, DevOps, platform, or operations-heavy engineering roles.
- Experience supporting production systems and participating in on-call rotations.
- Ability to debug live systems under pressure.
- Experience operating cloud infrastructure, with AWS preferred.
- Working knowledge of Kubernetes and containerized workloads.
- Infrastructure as Code experience with Terraform or a similar tool.
- Familiarity with monitoring and alerting tools such as Datadog and Prometheus.
- Scripting or automation experience with Python, Bash, or similar tools.
- Experience leading incidents or mentoring others during on-call is preferred.
- Experience in regulated or security-sensitive environments is preferred.
- Familiarity with production databases, queues, and caches is preferred.
- Interest in SLOs, error budgets, and capacity planning is preferred.
Benefits
- Equity from day one.
- Personal development budget.
- Ability to work from anywhere for one month.
- Dedicated wellness days and birthday off.
- Hybrid work environment with 3 days in the office.
Tech Stack
Categories
About Heidi
Heidi is building an AI Care Partner to expand clinical capacity by supporting every stage of care delivery. In addition to its AI scribe, Heidi has introduced Evidence, giving clinicians access to trusted medical research at the point of care, and Comms, enabling healthcare teams to coordinate patient communications. Heidi supports more than 2.7 million patient interactions each week in 110 languages from 190 countries. Founded in Melbourne, Australia, Heidi has raised $96.6M USD from global investors including Point72 Private Investments, Blackbird, Headline, Phoenix Court's growth fund Latitude, Possible Ventures and Archangel. Heidi aligns with leading international healthcare and privacy frameworks, including NHS requirements, GDPR, HIPAA, and the Australian Privacy Principles, and maintains enterprise-grade security certifications including ISO 27001, SOC 2 Type II, Cyber Essentials Plus, and ISO 42001. Learn more at heidihealth.com
