Firmus Technologies

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies
Apply
3 hours ago
Sydney, AustraliaSenior

Responsibilities

  • Operate, automate, and continuously improve GPU interconnect and high-performance network fabrics, including NVLink, NVSwitch, InfiniBand, and Spectrum-X.
  • Build guarded automation and remediation tooling for fabric faults, fault isolation, and fabric reconvergence.
  • Diagnose and tune performance across the interconnect stack, including collective communication, link-level behavior, RoCE, and congestion control.
  • Operate DPU-based host networking, including offload paths and driver and firmware compatibility.
  • Plan and execute firmware upgrades, fabric expansions, capacity changes, verification, staged rollouts, and rollback procedures.
  • Provide deep technical diagnosis for fabric and interconnect failures and drive permanent fixes to closure.
  • Lead engineering-level vendor escalations and reproduce faults to vendor standards.
  • Lead technical recovery during major fabric incidents, participate in the follow-the-sun on-call roster, mentor engineers, and maintain runbooks and operational documentation.

Requirements

  • 8+ years of experience, including substantial ownership of production network or interconnect infrastructure in a 24/7 environment.
  • Deep operational experience with high-performance GPU interconnect fabrics, including NVLink and NVSwitch topology and failure diagnosis.
  • Extensive experience with InfiniBand or RoCE-based Ethernet such as NVIDIA Spectrum-X, including routing internals and congestion control tuning.
  • Experience with DPU or SmartNIC-based host networking, offload paths, and driver and firmware coordination.
  • Experience planning and executing production fabric firmware upgrades and capacity expansions with defined rollback procedures.
  • Strong infrastructure automation and infrastructure-as-code experience using peer review and progressive rollout practices.
  • Practical scripting or programming experience for operational automation and tooling, such as Python, Go, or Bash.
  • Experience serving as a senior production escalation point, including major incident response, on-call participation, vendor escalation, and runbook creation.
  • Experience with large-scale distributed training or inference fabrics, NVIDIA rack-scale or multi-node GPU systems, and multi-tenant cloud, service provider, or colocation environments is preferred.
  • Knowledge of data centre and hardware fundamentals, including cabling, optics, firmware management, and hardware fault workflows, is preferred; relevant vendor certification is also preferred.

Benefits

  • Permanent full-time employment.
  • Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
  • Participation in a shared after-hours escalation roster for the role’s domain within a 24/7 operations function.
  • Work under broad direction with substantial autonomy and direct access to decision makers.

Tech Stack

Categories

DevOpsSite Reliability
Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me