Senior Platform Reliability Engineer
Firmus Technologies14 hours ago
Melbourne, AustraliaSenior
Responsibilities
- Operate the multi-tenant control and management plane, including tenancy, quota, and access controls.
- Operate exabyte-scale distributed, parallel, and high-performance storage and S3-compatible object storage to defined service levels.
- Operate the production GPU compute fleet, including node health, fault detection, firmware and driver baselines, remediation, and hardware replacement workflows.
- Build and maintain guarded automation and remediation tooling that enables self-healing operations.
- Diagnose and tune performance across operating systems, networks, and software-defined storage data paths.
- Deliver infrastructure changes as code and improve cluster validation, CI/CD automation, provisioning, and testing frameworks.
- Operate shared observability infrastructure for metrics, logs, traces, alerting, and retention.
- Manage production readiness reviews, monthly patching, urgent vulnerability remediation, and service-level adherence.
- Provide deep technical escalation, root-cause analysis, permanent-fix coordination, and vendor escalation.
- Lead technical recovery during major incidents, participate in the follow-the-sun on-call roster, mentor engineers, and document runbooks and procedures.
Requirements
- 8+ years of infrastructure, systems, or platform engineering experience, including substantial ownership of production storage and shared infrastructure services in a 24/7 environment.
- Extensive experience with scale-out, parallel, or enterprise storage such as VAST, WEKA, Ceph, Lustre, GPFS, or NetApp.
- Substantial experience operating virtualization platforms such as Proxmox, VMware, or KVM and building shared services to defined service levels.
- Expert-level Linux knowledge covering storage and filesystem internals, kernels, drivers, networking, memory, I/O, and performance analysis.
- Extensive experience operating observability infrastructure across metrics, logs, traces, and alerting, including tools such as Prometheus, Grafana, OpenTelemetry, Loki, or Elasticsearch.
- Strong infrastructure automation, infrastructure-as-code, and GitOps experience using tools such as OpenTofu, Terraform, Ansible, Argo CD, or CircleCI.
- Practical scripting or programming experience for operational automation using Python, Go, or Bash.
- Working competence with Kubernetes and the ability to diagnose issues at the boundary between shared services and container estates.
- Experience with major incident response, on-call work, post-incident reviews, vendor escalation, and executable runbook creation.
- Experience applying least privilege, secrets management, certificate-based access, audited privileged operations, and audit evidence practices.
- Preferred experience with GPU or HPC infrastructure, multi-tenant cloud or colocation environments, GPU-enabled Kubernetes, bare-metal provisioning, high-performance networking, policy-as-code, software supply-chain controls, and ISO 27001 or SOC 2.
- Bachelor's degree in computer science, engineering, or a related discipline, or an equivalent combination of relevant experience and training.
Benefits
- Permanent full-time employment.
- Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
- Participation in a shared after-hours escalation roster within a 24/7 operational function.
- Work with large-scale sustainable AI infrastructure and high-performance GPU systems.
Tech Stack
Categories
DevOpsSite Reliability
About Firmus Technologies
Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.