
Distributed Systems Engineer
Fluidstack2 months ago
Austin, TX, USA +3 moreSenior
Base Salary
$208k - $269k/yr
Responsibilities
- Own the fleet observability platform, including data pipelines, decoration and correlation, health checks, and telemetry from sites to devices and links.
- Define versioned APIs and contracts for production infrastructure and internal fleet management tools.
- Build and operate the Kubernetes-based production control plane for machine management, state inspection, and distributed command execution.
- Maintain fleet state as a source of truth across provisioning, operations, infrastructure management, and customer-facing platforms.
- Own SLOs, site lifecycle state, and the systemic resolution of incidents and reliability issues.
- Integrate new XPU generations and sites through ZTP, DHCP, DNS, and artifact management.
Requirements
- Treat operational toil as an automation problem and take ownership through ambiguity.
- Have experience designing APIs that remain effective at scale and shipping production services used by other teams.
- Be comfortable with on-call operations, incident leadership, postmortems, and fixing systemic causes.
- Be able to learn unfamiliar domains quickly and work effectively across languages with AI coding tools.
- Be familiar with LLM APIs, MCP servers, agentic frameworks, Claude Code, Cursor, or similar tools.
- Distributed systems, data pipeline engineering, time-series observability, API versioning, workflow orchestration, hardware telemetry, Go, Python, and Postgres are bonus qualifications.
Benefits
- Competitive total compensation package including salary and equity.
- Retirement or pension plan in line with local norms.
- Health, dental, and vision insurance.
- Generous paid time off in line with local norms.