
Site Reliability Engineer, Compute
Fluidstack2 months ago
Austin, TX, USA +3 moreSenior
Base Salary
$208k - $269k/yr
Responsibilities
- Own end-to-end health, reliability, scalability, and operations for a large GPU compute fleet across Kubernetes and bare-metal environments.
- Build metrics pipelines, alerting, observability, and unified fleet-health views.
- Automate compute failure workflows from detection and triage through parts management, RMA, and return to service.
- Design and expand GPU qualification platforms for burn-in, performance baselining, and new product introduction.
- Own Redfish and BMC tooling, including firmware-level telemetry, log collection, and low-level fleet access.
- Run incidents, write postmortems, and address systemic reliability causes.
- Build production automation and eliminate manual repair and operational toil.
Requirements
- Experience shipping production automation that other teams depend on, using any programming language and AI coding tools.
- Comfort reasoning about hardware, firmware, silicon-level failure modes, and software infrastructure.
- Ability to operate independently in ambiguous, high-urgency environments and rapidly learn unfamiliar domains.
- Willingness to participate in pager-based incident response and own incidents through systemic resolution.
- Fluency with AI tooling such as LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, or similar tools.
- Hardware lifecycle management and RMA automation experience is preferred.
- Experience with BMC/Redfish or IPMI tooling is preferred.
- Experience with GPU qualification or burn-in frameworks is preferred.
- Experience with workflow and orchestration engines such as Temporal or Cadence is preferred.
- Experience with metrics and alerting tools such as Prometheus or Grafana is preferred.
- Go or Python experience is preferred.
Benefits
- Competitive total compensation package including salary and equity.
- Retirement or pension plan in line with local norms.
- Health, dental, and vision insurance.
- Generous paid time off policy in line with local norms.
- Equity may be provided in the form of restricted stock units.
Tech Stack
Categories
DevOpsSite Reliability