Rubrik

Staff Software Engineer - Reliability

Rubrik
Apply
3 months ago
Palo Alto, CA, USAStaff+
H1B Sponsor

Base Salary

$218k - $328k/yr

Responsibilities

  • Set the architectural direction for Rubrik’s cloud platform and optimize backend infrastructure for performance, security, and multi-region scale.
  • Build and maintain internal tools, platform controllers, and automation frameworks that reduce operational toil.
  • Establish cross-organizational reliability standards, resilience practices, capacity safeguards, and compliance controls.
  • Define and enforce reliability governance using service indicators, objectives, error budgets, and telemetry-driven roadmaps.
  • Act as an Incident Commander for severe cloud outages and lead blameless post-mortems and systemic remediation.
  • Lead cost observability, infrastructure capacity forecasting, resource quota optimization, and vendor service-level management.
  • Set the technical direction for the Application-SRE team and resolve complex customer-impacting platform issues.
  • Mentor engineers, contribute to interview frameworks, and raise the organization’s technical standards.
  • Partner with engineering, Sales, Support, customers, and Product to convert field signals into reliability and product roadmap improvements.
  • Participate in on-call rotations.
  • Follow federal security and privacy policies, support vulnerability remediation, and collaborate with Information Security on system controls.

Requirements

  • U.S. citizenship and current residence in the continental United States are required.
  • Bachelor’s, master’s, or doctoral degree in Computer Science, Computer Engineering, or a closely related technical discipline is required.
  • A minimum of 8–12+ years of software engineering and production cloud infrastructure experience is required, including at least 5+ years in formal SRE, DevOps, or Platform engineering roles.
  • Hands-on programming expertise in Go, Python, or Java, with knowledge of concurrency, data structures, and test-driven software design patterns, is required.
  • Experience designing, deploying, analyzing, and auditing large-scale distributed systems, database topologies, and highly available public cloud environments is required.
  • Strong Unix/Linux systems knowledge, systems administration experience, and advanced L4/L7 networking expertise are required.
  • Experience partnering with Sales, Support, customers, and Product on escalations, proofs of concept, and field-to-product feedback is required.
  • Demonstrated technical leadership across architectural dependencies, multi-team projects, and major platform changes is required.
  • Preferred experience includes enterprise-scale Kubernetes deployments using GKE or EKS and large-scale MySQL environments.
  • Preferred experience includes infrastructure compliance with FedRAMP, SOC 2, ISO 27001, or CJIS.
  • Preferred experience includes Terraform or Pulumi module design, multi-tenant state isolation, and observability using Prometheus, Grafana, or OpenTelemetry.

Benefits

  • U.S.-based role requiring current U.S. citizenship and residence in the continental United States.
  • Participation in on-call rotations.
  • May require Tier 1, Tier 2 Public Trust, or CJIS background screening depending on government-data access.
  • Eligible for equity and company benefits.

Tech Stack

Categories

DevOpsSite Reliability
Rubrik

About Rubrik

1,001-5,000 employees

Rubrik (NYSE: RBRK), the Security and AI company, operates at the intersection of data protection, cyber resilience and enterprise AI acceleration. The Rubrik Security Cloud platform is designed to deliver robust cyber resilience and recovery including identity resilience to ensure continuous business operations, all on top of secure metadata and data lake. Follow @rubrikInc on X (formerly Twitter), YouTube, and Instagram.

Contact me