Principal Engineer, AI Cloud Software
Firmus Technologies8 hours ago
Singapore, SingaporeStaff+
Responsibilities
- Define GPU and host health criteria and service-readiness gates for customer and internal AI workloads.
- Create reusable golden dashboards, low-noise alerts, PromQL/LogQL queries, health checks, reference implementations, and runbooks.
- Develop and operationalize DCGM checks, NCCL and bandwidth tests, stress and burn-in tests, and validation jobs with clear pass/fail criteria.
- Isolate GPU, cooling, host, network, power-limit, and management-interface faults using combined host and BMC telemetry.
- Detect non-crashing failures such as ECC errors, NVLink retries, thermal slowdown, XID patterns, stragglers, and silent data corruption.
- Partner with commissioning, operations, infrastructure, telemetry owners, and customer teams on bring-up, break-fix, incidents, and technical sessions.
- Participate in GPU and host incident response, including server-level debugging, on-call rotation, and converting findings into reusable operational knowledge.
- Provide engineering leadership with clear assessments of fleet health and risk.
Requirements
- Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
- At least 7 years of experience in GPU, HPC, AI infrastructure, or closely related large-scale systems engineering environments.
- Experience owning health monitoring or diagnostics used by customers or internal teams.
- Production experience diagnosing GPU faults involving XID events, ECC or memory errors, NVLink issues, and power or thermal limits.
- Strong Linux and server fundamentals, including debugging across GPU, CPU, memory, PCIe, power, and cooling.
- Experience automating infrastructure analysis with Python or a similar language.
- Hands-on experience with DCGM, NCCL or collective tests, stress testing, and validation checks.
- Experience analyzing infrastructure telemetry and building trusted dashboards and alerts with PromQL, LogQL, Grafana, or comparable tools.
- Ability to structure operational health knowledge for use by AI assistants and operations teams.
- Willingness to participate in incident-response on-call and occasionally travel overseas.
- Clear written and verbal communication in English.
- Highly desirable: experience with large GPU systems, multiple sites or GPU generations, AI or GPU clouds, HPC centers, GPU or host failure prediction, and AI-assisted operational knowledge.
Benefits
- Full-time employment based in Singapore.
- Occasional overseas travel is required when the role needs it.
- Participation in an incident-response on-call rotation.
About Firmus Technologies
Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.