2 hours ago
Responsibilities
- Own the architectural evolution toward cell-based architecture and isolated failure domains.
- Define the Kubernetes strategy, including multi-cluster management, service mesh, automated scaling, and Golden Path developer experiences.
- Set organization-wide standards for SRE, causal observability, automated incident response, and SLO/SLI management.
- Lead the technical strategy and tooling for FinOps, including cost-per-service visibility and infrastructure cost optimization.
- Build agent-friendly developer infrastructure, including standardized Dev Containers and ephemeral environments.
- Guide multiple teams toward coherent long-term technical direction and make high-impact architectural decisions.
- Design, prototype, build, deploy, and operate reliable cloud-native platform services at scale.
Requirements
- Extensive professional experience designing, building, and operating large-scale cloud-native distributed systems and platform services on Kubernetes.
- Proven ownership of critical services or multi-service platforms, including low-level design, API design, data modeling, deployment, and operational health.
- Deep expertise with at least one major public cloud provider and core platform technologies covering compute, networking, storage, service discovery, security, observability, and CI/CD.
- Demonstrated ability to make high-impact architectural decisions, navigate complex trade-offs, and guide multiple teams.
- Familiarity with AI-driven systems, tools, or workflows and applying AI/ML concepts to real-world cloud or platform products.
- Deep knowledge of OpenTelemetry, Prometheus, and distributed tracing.
- Expert-level experience with Infrastructure as Code using Terraform or Pulumi and CI/CD at scale.
- Proficiency in Go, Rust, or similar languages used in modern platform engineering.
- Track record of defining multi-year technical strategies and driving adoption of shared cloud or developer platforms across many teams.
- Experience designing and operating highly available, globally distributed systems, including capacity planning, performance optimization, and failure handling.
- Experience safely integrating AI/ML-enabled solutions such as intelligent routing, predictive scaling, or automated remediation into platform services.
- Advanced experience applying AI/ML to cloud and platform problems and partnering with data/ML teams to productionize these capabilities.
- Experience with cloud architecture, failure domains, latency, unit economics, global-scale reliability, and hands-on production-grade infrastructure development.
Benefits
- Medical, dental, and vision coverage.
- Paid time off, an Employee Assistance Program, wellness reimbursement, and travel reimbursement.
- Travel discounts and International Airlines Travel Agent Network membership.
- The posting includes San Jose location-specific logistics and states that pay varies based on location, budget, knowledge, skills, and experience.
