Fluidstack

Software Engineer, Compute Operations

Fluidstack
Apply
5 hours ago

Base Salary

$224k - $300k/yr

Responsibilities

  • Build a real-time fleet health system for machines across Kubernetes and bare metal, including telemetry, tiered health checks, alarms, incident correlation, and probable-cause drafting.
  • Automate repair and RMA workflows from failure detection through triage, parts, vendor return, and return to service.
  • Develop rack-level workflows for hardware burn-in, performance baselining, and new hardware validation.
  • Own facility asset models, maintenance systems, work orders, data migration, and retirement of legacy datacenter inventory tools.
  • Structure SOPs, training records, technician qualifications, and operational metrics into auditable systems and dashboards.
  • Work forward-deployed with production engineers and facility operators, including on site and on the on-call rotation.

Requirements

  • Production experience shipping code in Go, Python, or TypeScript, with the ability to learn additional languages as needed.
  • Experience building production features with LLM APIs such as OpenAI or Anthropic, open-weight models, MCP servers, and agentic frameworks.
  • Daily use of AI coding tools such as Claude Code and Cursor, including autonomous agent workflows.
  • Ability to identify problems, design solutions, and ship independently under deadline pressure.
  • Experience participating in or working alongside an on-call rotation and turning operational pain into quieter, more effective systems.
  • Strong product judgment for building clear interfaces and workflows that match operational work.
  • Preferred experience includes production engineering or SRE for large GPU fleets, hardware qualification or burn-in frameworks, BMC, Redfish, IPMI, CMMS, DCIM, asset management systems, BMS, EPMS, SCADA, Prometheus, and Grafana.

Benefits

  • Competitive total compensation package including cash and equity.
  • Health, dental, and vision insurance.
  • Retirement plan.
  • Generous paid time off policy.

Tech Stack

Categories

Fluidstack

About Fluidstack

201-500 employees

Fluidstack provides on-demand GPU cloud infrastructure for AI training and inference, offering dedicated clusters of NVIDIA H200 and GB200 systems. The company sells compute capacity and managed clusters to AI labs, enterprises, and public-sector teams that need high-throughput workloads. Founded in 2017 and headquartered in New York, it is privately held.

Contact me