5 hours ago
Responsibilities
- Improve production reliability and system resilience within an SRE-scoped team.
- Participate in on-call rotations, incident response, and blameless post-incident reviews.
- Write code and tooling, handle alerts, automate operational work, and reduce toil.
- Create runbooks that support the broader team and prevent recurring problems.
- Operate and improve production observability across metrics, logs, and traces.
- Communicate with teams and stakeholders during requirements analysis, demonstrations, and technical work.
- Troubleshoot complex production infrastructure issues and support teammates.
- Champion DevOps, SRE culture, and industry best practices.
Requirements
- 5+ years administering Linux systems and related infrastructure in production environments.
- Familiarity with SRE concepts including SLIs, SLOs, SLAs, error budgets, blast radius, and blameless postmortems.
- Strong Kubernetes and broader Kubernetes ecosystem fundamentals.
- Cloud infrastructure experience; AWS is strongly preferred and bare-metal experience is a bonus.
- Strong tool development skills using Bash plus Python, Go, or similar.
- Experience with infrastructure-as-code tooling, with Terraform preferred.
- Experience with CI/CD and version control, with GitHub preferred.
- Database experience with Postgres, Cassandra, or ClickHouse preferred.
- Experience operating production observability stacks covering metrics, logs, and traces.
- Strong troubleshooting instincts and ownership of incident response.
- Evidence of continual professional development and ability to work independently in an async, globally distributed team.
Benefits
- Remote-first flexible working environment with coworking options.
- Four weeks of paid annual leave, parental leave, birthday leave, and a purchased annual leave program.
- Wellness allowance and employee wellbeing initiatives.
- Study and training allowance plus five days of paid study leave.
- Modern workspaces available when not working remotely.
- Inclusive team environment with industry experts and fresh talent.
- Legend and Kudos recognition programs.
Tech Stack
Categories
DevOpsSite Reliability
About Megaport
Megaport provides network-as-a-service via a software-defined network that lets enterprises provision private, on-demand connections among data centers, branch sites, and public clouds like AWS and Google Cloud. Founded in 2013 and headquartered in Fortitude Valley, Queensland, it is a public company listed on the ASX (MP1) and operates connectivity to 850+ data centers across 25+ countries. Revenue comes from subscription and usage-based interconnection and cloud on-ramp services.
