Firmus Technologies

Principal Engineer, AI Cloud Software

Firmus Technologies
Apply
8 hours ago
Singapore, SingaporeStaff+

Responsibilities

  • Define GPU and host health criteria and service-readiness gates for customer and internal AI workloads.
  • Create reusable golden dashboards, low-noise alerts, PromQL/LogQL queries, health checks, reference implementations, and runbooks.
  • Develop and operationalize DCGM checks, NCCL and bandwidth tests, stress and burn-in tests, and validation jobs with clear pass/fail criteria.
  • Isolate GPU, cooling, host, network, power-limit, and management-interface faults using combined host and BMC telemetry.
  • Detect non-crashing failures such as ECC errors, NVLink retries, thermal slowdown, XID patterns, stragglers, and silent data corruption.
  • Partner with commissioning, operations, infrastructure, telemetry owners, and customer teams on bring-up, break-fix, incidents, and technical sessions.
  • Participate in GPU and host incident response, including server-level debugging, on-call rotation, and converting findings into reusable operational knowledge.
  • Provide engineering leadership with clear assessments of fleet health and risk.

Requirements

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • At least 7 years of experience in GPU, HPC, AI infrastructure, or closely related large-scale systems engineering environments.
  • Experience owning health monitoring or diagnostics used by customers or internal teams.
  • Production experience diagnosing GPU faults involving XID events, ECC or memory errors, NVLink issues, and power or thermal limits.
  • Strong Linux and server fundamentals, including debugging across GPU, CPU, memory, PCIe, power, and cooling.
  • Experience automating infrastructure analysis with Python or a similar language.
  • Hands-on experience with DCGM, NCCL or collective tests, stress testing, and validation checks.
  • Experience analyzing infrastructure telemetry and building trusted dashboards and alerts with PromQL, LogQL, Grafana, or comparable tools.
  • Ability to structure operational health knowledge for use by AI assistants and operations teams.
  • Willingness to participate in incident-response on-call and occasionally travel overseas.
  • Clear written and verbal communication in English.
  • Highly desirable: experience with large GPU systems, multiple sites or GPU generations, AI or GPU clouds, HPC centers, GPU or host failure prediction, and AI-assisted operational knowledge.

Benefits

  • Full-time employment based in Singapore.
  • Occasional overseas travel is required when the role needs it.
  • Participation in an incident-response on-call rotation.

Tech Stack

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me