
Software Engineer, Compute (GPU)
Fluidstack2 months ago
Austin, TX, USA +3 moreSenior
Base Salary
$208k - $269k/yr
Responsibilities
- Own compute fleet health, metrics pipelines, alerting, and unified GPU health views across Kubernetes and bare metal.
- Automate compute failure detection, triage, parts management, repair, and return-to-service workflows.
- Design and expand GPU qualification platforms covering burn-in, performance baselining, and new product introduction execution.
- Build and maintain Redfish and BMC tooling for firmware-level telemetry, log collection, and low-level fleet access.
- Own the reliability, scalability, operations, and incident response discipline of the production compute fleet.
- Write postmortems and address systemic causes of incidents.
Requirements
- Experience shipping production automation that other teams depend on and comfort working in any programming language with AI coding tools.
- Ability to reason about hardware failure modes at the firmware and silicon level.
- Comfort carrying a pager, running incidents, writing postmortems, and resolving systemic causes.
- Ability to work independently in ambiguous domains and quickly develop competence in unfamiliar areas.
- Fluency with AI tooling, including LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, or similar tools.
- Preferred experience with hardware lifecycle management, RMA automation, BMC/Redfish or IPMI tooling, GPU qualification or burn-in frameworks, Temporal or Cadence, Prometheus or Grafana, and Go or Python.
Benefits
- Competitive total compensation package including salary and equity.
- Retirement or pension plan in line with local norms.
- Health, dental, and vision insurance.
- Generous paid time off policy in line with local norms.