Fluidstack

Software Engineer, Compute (GPU)

Fluidstack
Apply
2 months ago
Austin, TX, USA +3 moreSenior

Base Salary

$208k - $269k/yr

Responsibilities

  • Own compute fleet health, metrics pipelines, alerting, and unified GPU health views across Kubernetes and bare metal.
  • Automate compute failure detection, triage, parts management, repair, and return-to-service workflows.
  • Design and expand GPU qualification platforms covering burn-in, performance baselining, and new product introduction execution.
  • Build and maintain Redfish and BMC tooling for firmware-level telemetry, log collection, and low-level fleet access.
  • Own the reliability, scalability, operations, and incident response discipline of the production compute fleet.
  • Write postmortems and address systemic causes of incidents.

Requirements

  • Experience shipping production automation that other teams depend on and comfort working in any programming language with AI coding tools.
  • Ability to reason about hardware failure modes at the firmware and silicon level.
  • Comfort carrying a pager, running incidents, writing postmortems, and resolving systemic causes.
  • Ability to work independently in ambiguous domains and quickly develop competence in unfamiliar areas.
  • Fluency with AI tooling, including LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, or similar tools.
  • Preferred experience with hardware lifecycle management, RMA automation, BMC/Redfish or IPMI tooling, GPU qualification or burn-in frameworks, Temporal or Cadence, Prometheus or Grafana, and Go or Python.

Benefits

  • Competitive total compensation package including salary and equity.
  • Retirement or pension plan in line with local norms.
  • Health, dental, and vision insurance.
  • Generous paid time off policy in line with local norms.

Tech Stack

GoGrafanaKubernetesPrometheusPython

Categories

DevOpsEmbeddedSite Reliability
Fluidstack

About Fluidstack

201-500 employees
Contact me