Backblaze External Website

Site Reliability Engineer ll (DBA)

Backblaze External Website
Apply
3 hours ago
Remote, Costa Rica +3 moreMid Level
H1B Sponsor

Responsibilities

  • Operate and maintain high-availability Vitess/MySQL and Cassandra database systems using established architecture and runbooks.
  • Optimize database performance through query tuning, indexing strategies, and schema design.
  • Execute documented backup, recovery, replication, security, and access-control procedures.
  • Monitor production service health using SLIs, SLOs, and error budgets, and escalate issues when thresholds are at risk.
  • Participate in on-call rotations, incident response, post-incident reviews, capacity planning, and disaster recovery exercises.
  • Develop automation and scripts to reduce operational toil and improve system reliability.
  • Contribute to monitoring, logging, alerting, CI/CD, configuration management, and infrastructure as code tooling.
  • Collaborate with engineering, product, operations, vendors, and service providers to troubleshoot issues and improve resilient operations.
  • Maintain playbooks, runbooks, operational documentation, and reliability-focused processes.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
  • 2–4 years of experience in site reliability, systems engineering, or operations centered around database systems.
  • Experience with large-scale, production-grade systems and established database operations procedures.
  • Solid Linux systems administration and troubleshooting skills.
  • Proficiency in at least one of Python, Bash, or Go.
  • Familiarity with monitoring, alerting, incident response, and root cause analysis.
  • Understanding of Kubernetes and Docker, with Kubernetes/Vitess-on-Kubernetes experience preferred.
  • Hands-on MySQL performance tuning, replication, and disaster recovery experience; Vitess or distributed/sharded MySQL experience is a plus.
  • Proficiency in SQL and NoSQL database management.
  • Experience in a SaaS, service provider, or distributed systems environment is preferred.
  • Familiarity with ITIL/OSS practices and SLOs/SLAs is preferred.
  • Experience with AWS, GCP, or Azure is preferred.
  • Ability to work independently, take ownership, and drive projects from discovery through resolution.

Tech Stack

AnsibleApache CassandraAWSAzureBashDockerGoGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxMySQLPrometheusPythonSQLTerraform

Categories

Site Reliability
Backblaze External Website

About Backblaze External Website

201-500 employees

AI infrastructure runs on data. Storage gets hard at scale. We’ve been breaking, rebuilding, and improving that storage layer for 20 years. We started by building storage servers from scratch in a Palo Alto apartment to save 90% on hardware. Turns out hard constraints make for pretty good engineering. That discipline forged into a platform that performs as well or better than competitors, scales to multi-exabyte workloads, and costs a fraction of what AWS or Google charge. No blank checks or architectural sprawl. Today, the world's leading AI infrastructure providers, neoclouds, and data-intensive businesses run on Backblaze. Our platform includes: - B2 Overdrive: High-throughput object storage engineered for AI infrastructure, GPU workloads, and massive-scale data pipelines. - B2 Neo: White-label enterprise cloud storage powering neocloud platforms for AI, HPC, and media. - B2 Cloud Storage: S3-compatible object storage with predictable pricing, free egress, and no lock-in. - Computer Backup: Automatic, continuous endpoint backup for Macs and PCs. 500,000+ customers. 175 countries. Nasdaq: BLZE. backblaze.com

Contact me