
Site Reliability Engineer II
Backblaze External WebsiteResponsibilities
- Support the availability and durability of critical services across production environments.
- Monitor service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk.
- Participate in on-call rotations, incident response, and post-incident reviews.
- Develop automation to reduce manual operational work and toil.
- Contribute to monitoring, logging, alerting, CI/CD, configuration management, and infrastructure as code tooling.
- Partner with engineering, product, and operations teams on resilient system design and operations.
- Assist with capacity planning, disaster recovery exercises, vendor troubleshooting, and SLA performance tracking.
- Create playbooks, runbooks, and operational documentation while identifying long-term reliability improvements.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 2–4 years of experience in site reliability, systems engineering, or operations.
- Exposure to large-scale, production-grade systems.
- Solid Linux systems administration and troubleshooting skills.
- Proficiency in at least one of Python, Bash, or Go.
- Understanding of monitoring, alerting, incident response, root cause analysis, containers, and microservices concepts.
- Preferred experience in a SaaS, service provider, or distributed systems environment.
- Preferred familiarity with ITIL/OSS practices, SLOs, SLAs, and cloud platforms such as AWS, GCP, or Azure.
Tech Stack
Categories
About Backblaze External Website
AI infrastructure runs on data. Storage gets hard at scale. We’ve been breaking, rebuilding, and improving that storage layer for 20 years. We started by building storage servers from scratch in a Palo Alto apartment to save 90% on hardware. Turns out hard constraints make for pretty good engineering. That discipline forged into a platform that performs as well or better than competitors, scales to multi-exabyte workloads, and costs a fraction of what AWS or Google charge. No blank checks or architectural sprawl. Today, the world's leading AI infrastructure providers, neoclouds, and data-intensive businesses run on Backblaze. Our platform includes: - B2 Overdrive: High-throughput object storage engineered for AI infrastructure, GPU workloads, and massive-scale data pipelines. - B2 Neo: White-label enterprise cloud storage powering neocloud platforms for AI, HPC, and media. - B2 Cloud Storage: S3-compatible object storage with predictable pricing, free egress, and no lock-in. - Computer Backup: Automatic, continuous endpoint backup for Macs and PCs. 500,000+ customers. 175 countries. Nasdaq: BLZE. backblaze.com