
Product Engineer, Compute Operations
Fluidstack6 hours ago
Base Salary
$224k - $300k/yr
Responsibilities
- Build a real-time fleet health system for machines across Kubernetes and bare metal, including telemetry, tiered health checks, alarms, incident correlation, and probable-cause reporting.
- Automate repair and RMA workflows from failure detection through triage, parts, vendor return, and return to service.
- Build software workflows for hardware burn-in, performance baselining, and validation of new hardware at rack scale.
- Own facility maintenance systems, asset models, data migration, deployment across sites, and retirement of legacy data-center inventory tools.
- Convert runbooks, SOPs, training records, and technician qualifications into structured, auditable data and build dashboards for site SLOs, deployment cycle time, and labor ramp.
- Work forward-deployed with production engineers and facility operators, including on-site work and on-call rotation participation.
Requirements
- Production experience shipping software in Go, Python, or TypeScript, with the ability to learn additional languages as needed.
- Experience building real features with LLM APIs such as OpenAI, Anthropic, or open-weight models, as well as MCP servers and agentic frameworks.
- Daily experience using AI coding tools such as Claude Code and Cursor, including autonomous agent workflows.
- Ability to identify problems, design solutions, and deliver independently under demanding deadlines.
- Experience working on an on-call rotation or alongside on-call teams and turning operational issues into systems improvements.
- Strong product judgment for building intuitive interfaces and workflows that match real operational practices.
- Bonus experience with production engineering or SRE for large GPU fleets, hardware qualification or burn-in frameworks, BMC, Redfish, IPMI, CMMS, DCIM, asset management systems, BMS, EPMS, SCADA, Prometheus, or Grafana.
Benefits
- Competitive total compensation package including cash and equity.
- Health, dental, and vision insurance.
- Retirement plan.
- Generous paid time off policy.
Tech Stack
Categories
Product Engineering
About Fluidstack
Fluidstack provides on-demand GPU cloud infrastructure for AI training and inference, offering dedicated clusters of NVIDIA H200 and GB200 systems. The company sells compute capacity and managed clusters to AI labs, enterprises, and public-sector teams that need high-throughput workloads. Founded in 2017 and headquartered in New York, it is privately held.