Senior AI Infrastructure Engineer, Observability
Firmus Technologies1 day ago
Singapore, SingaporeSenior
Responsibilities
- Define GPU and host health criteria and service-readiness gates for customer and internal AI workloads.
- Create reusable golden dashboards, alerts, PromQL/LogQL queries, health checks, and operational guidance.
- Develop DCGM checks, NCCL and bandwidth tests, stress and burn-in tests, and validation jobs with clear pass/fail criteria.
- Diagnose and isolate GPU, host, network, cooling, power, and management-interface faults using combined telemetry.
- Detect degradation and silent failures such as ECC errors, NVLink retries, thermal slowdown, XID patterns, and incorrect results.
- Publish reference implementations, runbooks, and structured operational knowledge for engineering and operations teams.
- Collaborate with commissioning, operations, infrastructure, telemetry owners, customer teams, and engineering leadership.
- Participate in GPU and host incident response and the on-call rotation.
Requirements
- Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
- 7+ years of experience in GPU, HPC, AI infrastructure, or closely related large-scale systems engineering environments.
- Experience owning health monitoring or diagnostics used by customers or internal teams.
- Production experience diagnosing GPU faults involving XID events, ECC or memory errors, NVLink issues, and power or thermal limits.
- Strong Linux and server fundamentals, including debugging across GPU, CPU, memory, PCIe, power, and cooling systems.
- Ability to automate with Python or a similar language.
- Hands-on experience with DCGM, NCCL or collective tests, stress testing, and reusable validation checks.
- Experience analyzing infrastructure telemetry and building trusted dashboards and alerts using PromQL, LogQL, Grafana, or comparable tooling.
- Ability to structure health knowledge for use by AI assistants during operational recovery.
- Willingness to participate in incident-response on-call and travel overseas occasionally.
- Clear written and verbal communication in English.
- Highly desirable: experience with large GPU systems, multiple sites or GPU generations, AI clouds, GPU clouds, HPC centers, failure prediction, or structured operational knowledge for AI-assisted incident recovery.
Benefits
- Full-time employment based in Singapore.
- Occasional overseas travel is required when the role requires it.
- Participation in an incident-response on-call rotation.
- Inclusive workplace committed to diversity and sustainability.