Firmus Technologies

Infrastructure Automation Engineer, AI Cluster Commissioning

Firmus Technologies
Apply
15 hours ago
Melbourne, AustraliaSenior

Responsibilities

  • Develop custom tools and automate hardware bring-up, configuration, firmware updates, inventory, testing, and issue tracking.
  • Implement and adapt monitoring, health-check, diagnostics, and remediation tools for new platforms and production integration.
  • Automate network deployment, firmware updates, configuration, diagnostics, and monitoring from the physical layer through the application layer.
  • Develop automation frameworks for storage deployment, configuration, benchmarking, and validation.
  • Integrate messaging, inventory, issue-tracking, reporting, dashboard, and operational analytics tools.
  • Develop and execute customized testing and benchmarking tools for project and customer requirements.
  • Use infrastructure-as-code tools and CI/CD pipelines to automate commissioning workflows and improve consistency and auditability.
  • Apply secure practices for updates, credentials, keys, secrets, system setup, network configuration, and storage configuration.
  • Document tools and processes and prepare them for operational handover.

Requirements

  • Bachelor’s degree in computer science, engineering, or a related technical field.
  • At least 5 years of experience developing software, automation platforms, operational tooling, or infrastructure automation solutions for large-scale Linux, cloud, HPC, or AI environments.
  • Strong understanding of high-performance GPU infrastructure, high-end networking, and high-performance storage.
  • Experience configuring Linux servers, networks, and storage systems.
  • Experience developing production software tools and APIs, integrating APIs, applying modern software development practices, and building automated testing frameworks.
  • Experience with Bash, Python, and Ansible for scripting and automation.
  • Experience with Redfish, IPMI, DCGM, Prometheus, Grafana, and NetBox.
  • Familiarity with Slurm, NVIDIA DCGM, NCCL, Kubernetes, and distributed-system validation frameworks.
  • Strong problem-solving, analytical, written, verbal communication, teamwork, and independent-working skills.
  • Willingness to travel internationally and domestically for on-site deployments and commissioning.

Benefits

  • Full-time employment based in Australia or Singapore.
  • Regular visits to current and future project sites in Australia and Southeast Asia.
  • Opportunity to work with founders and experts on large-scale AI infrastructure, energy systems, and next-generation compute.

Tech Stack

AnsibleBashGrafanaKubernetesLinuxPrometheusPython

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me