Senior Platform Reliability Engineer (Fabric and Interconnect)
Firmus Technologies3 hours ago
Sydney, AustraliaSenior
Responsibilities
- Operate, automate, and continuously improve GPU interconnect and high-performance network fabrics, including NVLink, NVSwitch, InfiniBand, and Spectrum-X.
- Build guarded automation and remediation tooling for fabric faults, fault isolation, and fabric reconvergence.
- Diagnose and tune performance across the interconnect stack, including collective communication, link-level behavior, RoCE, and congestion control.
- Operate DPU-based host networking, including offload paths and driver and firmware compatibility.
- Plan and execute firmware upgrades, fabric expansions, capacity changes, verification, staged rollouts, and rollback procedures.
- Provide deep technical diagnosis for fabric and interconnect failures and drive permanent fixes to closure.
- Lead engineering-level vendor escalations and reproduce faults to vendor standards.
- Lead technical recovery during major fabric incidents, participate in the follow-the-sun on-call roster, mentor engineers, and maintain runbooks and operational documentation.
Requirements
- 8+ years of experience, including substantial ownership of production network or interconnect infrastructure in a 24/7 environment.
- Deep operational experience with high-performance GPU interconnect fabrics, including NVLink and NVSwitch topology and failure diagnosis.
- Extensive experience with InfiniBand or RoCE-based Ethernet such as NVIDIA Spectrum-X, including routing internals and congestion control tuning.
- Experience with DPU or SmartNIC-based host networking, offload paths, and driver and firmware coordination.
- Experience planning and executing production fabric firmware upgrades and capacity expansions with defined rollback procedures.
- Strong infrastructure automation and infrastructure-as-code experience using peer review and progressive rollout practices.
- Practical scripting or programming experience for operational automation and tooling, such as Python, Go, or Bash.
- Experience serving as a senior production escalation point, including major incident response, on-call participation, vendor escalation, and runbook creation.
- Experience with large-scale distributed training or inference fabrics, NVIDIA rack-scale or multi-node GPU systems, and multi-tenant cloud, service provider, or colocation environments is preferred.
- Knowledge of data centre and hardware fundamentals, including cabling, optics, firmware management, and hardware fault workflows, is preferred; relevant vendor certification is also preferred.
Benefits
- Permanent full-time employment.
- Based in Australia or Singapore, with travel to Australian AI Factory sites as required.
- Participation in a shared after-hours escalation roster for the role’s domain within a 24/7 operations function.
- Work under broad direction with substantial autonomy and direct access to decision makers.
About Firmus Technologies
Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.