
Software Engineer, GPU Infrastructure
Fluidstack2 months ago
Austin, TX, USA +3 moreSenior
Base Salary
$208k - $269k/yr
Responsibilities
- Own compute fleet health end to end through metrics pipelines, alerting, and unified health views across Kubernetes-orchestrated workloads and bare metal.
- Build automation for compute failure detection, triage, parts management, repair, and return to service.
- Design and expand GPU qualification platforms covering burn-in, performance baselining, and new product introduction execution.
- Own Redfish and BMC tooling for firmware-level telemetry, log collection, and low-level fleet access.
- Drive reliability, scalability, observability, incident response, and operational automation for a hyperscale GPU fleet.
- Carry a pager, run incidents, write postmortems, and fix systemic causes.
- Use AI tooling and coding assistants to build production automation in the language best suited to the problem.
Requirements
- Experience shipping production automation that other teams depend on.
- Comfort reasoning about hardware failure modes at the firmware and silicon level.
- Ability to operate effectively in ambiguous, unfamiliar technical domains and learn quickly.
- Willingness to participate in on-call incident response and postmortem-driven improvement.
- Fluency with AI tooling, including LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, or similar tools.
- Bonus experience with hardware lifecycle management, RMA automation, BMC or Redfish tooling, GPU qualification or burn-in frameworks, workflow orchestration engines, metrics and alerting pipelines, Go, or Python.
Benefits
- Competitive total compensation package including salary and equity.
- Retirement or pension plan in line with local norms.
- Health, dental, and vision insurance.
- Generous paid time off policy in line with local norms.
- Equity may be provided in the form of restricted stock units.
- The posting states a commitment to pay equity and transparency.
Tech Stack
Categories
DevOpsSite Reliability