8 months ago
Responsibilities
- Improve the reliability, availability, and operational health of production systems.
- Set observability standards for metrics, logs, and errors across services.
- Define SLOs and SLIs, configure alerting, and improve on-call readiness and signal quality.
- Partner with engineers to design resilient systems and reduce operational risk early.
- Build internal tooling for system safety, debugging, and developer velocity.
- Manage infrastructure through Pulumi across GCP, AWS, and Firebase.
Requirements
- At least 5 years of SRE, DevOps, or production operations experience is required.
- At least 2 years of TypeScript web application development experience is required.
- Experience operating and scaling production systems against uptime and latency goals is required.
- Hands-on experience with observability stacks such as Datadog or Sentry is required.
- Experience defining SLOs and SLIs, building alerting strategies, and preparing systems for on-call operations is required.
- Proficiency with CI/CD systems and infrastructure-as-code is required.
- Experience with cloud-native and serverless platforms, including GCP and AWS, is required.
- Strong cross-system debugging and incident response skills are required.
- Preferred qualifications include distributed cloud tracing, Cloud Run, Firebase, Lambda, large TypeScript monorepos, PNPM, AI-powered systems, and startup or rapid-growth experience.
Tech Stack
Categories
DevOpsSite Reliability
About Flux
Flux is a better way to build PCBs. Now with Copilot, your AI assistant! 🚀 Sign up for free at http://flux.ai!
