
Member of Technical Staff - Reliability Engineering
Fireworks AI2 months ago
Base Salary
$240k - $290k/yr
Responsibilities
- Define reliability standards including SLOs, error budgets, and production-readiness criteria using real system telemetry.
- Own logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
- Identify and resolve customer-facing reliability issues that may span multiple services or fall outside individual service SLOs.
- Investigate cross-system failure modes involving retries, timeouts, and undocumented dependencies, and drive fixes with responsible teams.
- Coordinate live production incidents, conduct blameless postmortems, and track corrective actions to completion.
- Automate repetitive operational work to reduce toil and on-call burden.
- Partner with cloud infrastructure, inference, training, performance, product, and control-plane teams on capacity, failover, rollouts, and customer reliability.
Requirements
- 5+ years of experience with Linux internals, system performance troubleshooting, and networking fundamentals including TCP/IP, HTTP, and gRPC.
- 5+ years of production software engineering in Python, Go, C++, or Rust.
- Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
- Experience with distributed systems, high-throughput control planes, microservices, or multi-region deployments.
- Knowledge of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
- Bachelor’s or Master’s degree in Computer Science or Computer Engineering, or equivalent practical experience.
- Ability to influence teams and drive adoption of reliability standards without direct authority.
- Preferred experience with Prometheus, Grafana, OpenTelemetry, GPU and ML infrastructure, AI-assisted operations, or open-source infrastructure and serving projects.
Benefits
- Work focused on AI infrastructure and model serving at a fast-growing Series D company.
- Collaborate with engineers and AI researchers on high-impact systems problems.
- Inclusive, equal-opportunity workplace with an emphasis on ownership, learning, and teamwork.
Categories
Site Reliability
About Fireworks AI
Fireworks AI builds a generative AI platform for developers and enterprises to train, fine-tune, and serve open models for production use across text, image, audio, embeddings, and multimodal workloads. It offers managed inference and tooling via APIs on globally distributed infrastructure, with a usage-based SaaS model. Founded in 2022 and headquartered in San Mateo, CA, Fireworks AI is a privately held, Series D company.