
Member of Technical Staff - Reliability Engineering
Fireworks AI3 hours ago
Base Salary
$240k - $290k/yr
Responsibilities
- Define reliability standards including SLOs, error budgets, and production-readiness criteria using real system telemetry.
- Own logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
- Identify and resolve customer-facing reliability issues that may span multiple services or fall outside individual service SLOs.
- Investigate cross-system failure modes involving retries, timeouts, and undocumented dependencies, and drive fixes with responsible teams.
- Coordinate live production incidents, conduct blameless postmortems, and track corrective actions to completion.
- Automate repetitive operational work to reduce toil and on-call burden.
- Partner with cloud infrastructure, inference, training, performance, product, and control-plane teams on capacity, failover, rollouts, and customer reliability.
Requirements
- 5+ years of experience with Linux internals, system performance troubleshooting, and networking fundamentals including TCP/IP, HTTP, and gRPC.
- 5+ years of production software engineering in Python, Go, C++, or Rust.
- Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
- Experience with distributed systems, high-throughput control planes, microservices, or multi-region deployments.
- Knowledge of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
- Bachelor’s or Master’s degree in Computer Science or Computer Engineering, or equivalent practical experience.
- Ability to influence teams and drive adoption of reliability standards without direct authority.
- Preferred experience with Prometheus, Grafana, OpenTelemetry, GPU and ML infrastructure, AI-assisted operations, or open-source infrastructure and serving projects.
Benefits
- Work focused on AI infrastructure and model serving at a fast-growing Series D company.
- Collaborate with engineers and AI researchers on high-impact systems problems.
- Inclusive, equal-opportunity workplace with an emphasis on ownership, learning, and teamwork.
Categories
Site Reliability
About Fireworks AI
Fireworks is the fastest way to build, tune, and scale AI on open models. Ship production-ready AI in seconds on our globally distributed cloud infrastructure, optimized for your use case. Fireworks powers production workloads at companies like Uber, Doordash, Notion, and Cursor—delivering 15× faster speed, 4× lower latency, and 4× more concurrency than closed models.