Firmus Technologies

Site Reliability Engineer, AI Infrastructure

Firmus Technologies
Apply
9 hours ago
Melbourne, AustraliaSenior

Responsibilities

  • Deploy, configure, and maintain GPU servers, storage servers, networking equipment, and software components.
  • Perform hardware diagnostics, systems-functionality checks, and firmware updates.
  • Assist with customer deployments involving bare-metal systems, HPC clusters, Kubernetes, and Slurm.
  • Troubleshoot hardware, network, software, and firmware-compliance issues as first-line engineering support.
  • Investigate incidents, escalate critical issues, and participate in a 24/7 on-call rotation.
  • Support the Global Operations Centre team with compute-infrastructure troubleshooting.
  • Document incidents, resolutions, lessons learned, and operational procedures.
  • Communicate with operations, engineering, internal teams, stakeholders, and end users on issue resolution.

Requirements

  • Bachelor’s degree in computer engineering, computer science, or a related technical field.
  • At least 5 years of experience in field service technical areas.
  • Strong knowledge of server hardware, firmware lifecycles, Linux environments, hardware troubleshooting, and physical and system-level security standards.
  • Experience with scripting languages such as Bash and Python.
  • Familiarity with configuration-management tools, CI/CD tools, workload managers, cluster software, and observability tools such as Slurm, Kubernetes, NVIDIA BCM, Prometheus, Grafana, and ELK.
  • Strong problem-solving, analytical, communication, and collaboration skills.

Benefits

  • Full-time employment.
  • Work location: Brooklyn, Melbourne, Australia.
  • Opportunity to work on AI infrastructure, GPU cloud, HPC systems, and energy-efficient computing alongside engineering and operations experts.

Tech Stack

BashGrafanaKubernetesLinuxPrometheusPython

Categories

DevOpsSite Reliability
Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me