Firmus Technologies

Senior AI Infrastructure Engineer, Observability

Firmus Technologies
Apply
1 day ago
Singapore, SingaporeSenior

Responsibilities

  • Define GPU and host health criteria and service-readiness gates for customer and internal AI workloads.
  • Create reusable golden dashboards, alerts, PromQL/LogQL queries, health checks, and operational guidance.
  • Develop DCGM checks, NCCL and bandwidth tests, stress and burn-in tests, and validation jobs with clear pass/fail criteria.
  • Diagnose and isolate GPU, host, network, cooling, power, and management-interface faults using combined telemetry.
  • Detect degradation and silent failures such as ECC errors, NVLink retries, thermal slowdown, XID patterns, and incorrect results.
  • Publish reference implementations, runbooks, and structured operational knowledge for engineering and operations teams.
  • Collaborate with commissioning, operations, infrastructure, telemetry owners, customer teams, and engineering leadership.
  • Participate in GPU and host incident response and the on-call rotation.

Requirements

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • 7+ years of experience in GPU, HPC, AI infrastructure, or closely related large-scale systems engineering environments.
  • Experience owning health monitoring or diagnostics used by customers or internal teams.
  • Production experience diagnosing GPU faults involving XID events, ECC or memory errors, NVLink issues, and power or thermal limits.
  • Strong Linux and server fundamentals, including debugging across GPU, CPU, memory, PCIe, power, and cooling systems.
  • Ability to automate with Python or a similar language.
  • Hands-on experience with DCGM, NCCL or collective tests, stress testing, and reusable validation checks.
  • Experience analyzing infrastructure telemetry and building trusted dashboards and alerts using PromQL, LogQL, Grafana, or comparable tooling.
  • Ability to structure health knowledge for use by AI assistants during operational recovery.
  • Willingness to participate in incident-response on-call and travel overseas occasionally.
  • Clear written and verbal communication in English.
  • Highly desirable: experience with large GPU systems, multiple sites or GPU generations, AI clouds, GPU clouds, HPC centers, failure prediction, or structured operational knowledge for AI-assisted incident recovery.

Benefits

  • Full-time employment based in Singapore.
  • Occasional overseas travel is required when the role requires it.
  • Participation in an incident-response on-call rotation.
  • Inclusive workplace committed to diversity and sustainability.

Tech Stack

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees
Contact me