Fireworks AI

Member of Technical Staff - Reliability Engineering

Fireworks AI
Apply
3 hours ago
San Mateo, CA, USA or New York, NY, USASenior
H1B Sponsor

Base Salary

$240k - $290k/yr

Responsibilities

  • Define reliability standards including SLOs, error budgets, and production-readiness criteria using real system telemetry.
  • Own logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling.
  • Identify and resolve customer-facing reliability issues that may span multiple services or fall outside individual service SLOs.
  • Investigate cross-system failure modes involving retries, timeouts, and undocumented dependencies, and drive fixes with responsible teams.
  • Coordinate live production incidents, conduct blameless postmortems, and track corrective actions to completion.
  • Automate repetitive operational work to reduce toil and on-call burden.
  • Partner with cloud infrastructure, inference, training, performance, product, and control-plane teams on capacity, failover, rollouts, and customer reliability.

Requirements

  • 5+ years of experience with Linux internals, system performance troubleshooting, and networking fundamentals including TCP/IP, HTTP, and gRPC.
  • 5+ years of production software engineering in Python, Go, C++, or Rust.
  • Experience operating and debugging Kubernetes, Terraform, and Docker in high-throughput production environments.
  • Experience with distributed systems, high-throughput control planes, microservices, or multi-region deployments.
  • Knowledge of fault-tolerant design, SLO/SLA management, automated failover, and high-availability architecture.
  • Bachelor’s or Master’s degree in Computer Science or Computer Engineering, or equivalent practical experience.
  • Ability to influence teams and drive adoption of reliability standards without direct authority.
  • Preferred experience with Prometheus, Grafana, OpenTelemetry, GPU and ML infrastructure, AI-assisted operations, or open-source infrastructure and serving projects.

Benefits

  • Work focused on AI infrastructure and model serving at a fast-growing Series D company.
  • Collaborate with engineers and AI researchers on high-impact systems problems.
  • Inclusive, equal-opportunity workplace with an emphasis on ownership, learning, and teamwork.

Categories

Site Reliability
Fireworks AI

About Fireworks AI

51-200 employees

Fireworks is the fastest way to build, tune, and scale AI on open models. Ship production-ready AI in seconds on our globally distributed cloud infrastructure, optimized for your use case. Fireworks powers production workloads at companies like Uber, Doordash, Notion, and Cursor—delivering 15× faster speed, 4× lower latency, and 4× more concurrency than closed models.