
Senior Site Reliability Engineer
Hyperbolic Labs5 months ago
Responsibilities
- Define and maintain service level objectives for job success rates and customer trust.
- Build incident response systems and lead incident response and post-mortem processes.
- Manage capacity, resource allocation, and cost efficiency across a distributed global GPU network.
- Design monitoring and alerting systems that provide deep infrastructure visibility.
- Build automation for capacity management and resource allocation.
- Implement secure progressive rollouts, canary deployments, and automated rollback mechanisms.
- Improve system resilience in collaboration with engineering teams.
- Implement tenant and workload isolation, network segmentation, security hardening, key management, and secure credential rotation.
- Build compliance frameworks for the cloud platform and maintain 24/7 operational reliability.
Requirements
- Expertise in site reliability engineering, including defining, monitoring, and maintaining SLOs and SLAs for production systems.
- Strong experience with capacity planning, resource allocation, forecasting, and cost optimization for distributed systems.
- Experience with incident response, on-call rotations, post-mortems, MTTR reduction, and system resilience improvements.
- Deep knowledge of progressive rollouts, canary deployments, feature flags, and automated rollback mechanisms.
- Proficiency with observability practices and tools such as metrics, logging, tracing, alerting systems, Prometheus, Grafana, and the ELK stack.
- Strong understanding of tenant isolation, workload isolation, network segmentation, and infrastructure security hardening.
- Experience with secrets management, key management systems, certificate management, and secure credential rotation.
- Knowledge of cloud security compliance frameworks including SOC 2 and ISO 27001.
- Experience with automation, infrastructure-as-code, configuration management, and CI/CD pipelines.
- Preferred: experience operating GPU infrastructure, AI/ML platforms, or compute marketplaces at scale.
- Preferred: background in distributed systems, peer-to-peer networks, or decentralized infrastructure.
- Preferred: knowledge of multi-tenancy security patterns, container security, runtime security tools, chaos engineering, fault injection, and resilience testing.
- Preferred: experience with cloud and GPU cost optimization and systems requiring 99.9%+ uptime SLAs.
- Preferred: background at AWS, Google Cloud, Azure, or infrastructure startups, or contributions to open-source reliability, observability, or security tools.
About Hyperbolic Labs
Hyperbolic provides high‑performance GPU clusters and managed inference for AI startups and ML teams that need reliable capacity on demand, at a lower cost than anyone else. Researchers use Hyperbolic to launch new models faster, avoid GPU waitlists, and scale from prototype to production on the same platform for both training and inference.