Levi Strauss & Co

Staff Site Reliability Engineer

Levi Strauss & Co
Apply
17 hours ago
Remote, SpainStaff+

Responsibilities

  • Define and operate SLOs, SLIs, and error budgets across platform services.
  • Improve MTTD and MTTR through observability, automated alerting, runbooks, incident response, and blameless post-mortems.
  • Identify and eliminate operational toil while building self-service infrastructure capabilities.
  • Automate deployments, configuration management, and operational workflows with Terraform, Helm, and GitOps.
  • Architect and optimize workloads across GCP and Azure, including self-healing infrastructure, autoscaling, and capacity planning.
  • Apply SRE practices to agentic AI and multi-agent systems, including latency objectives, fallback patterns, monitoring, drift detection, and rollback capabilities.
  • Champion security, governance, encryption, IAM, secrets management, audit logging, and data quality practices.
  • Mentor junior and mid-level SREs and lead reliability reviews, code reviews, and cross-functional reliability initiatives.

Requirements

  • Master's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 10+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in large-scale production environments.
  • Deep hands-on expertise with GCP services including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, Composer, Dataflow, and Vertex AI.
  • Proficiency with Terraform, Helm, and GitOps workflows using ArgoCD or Flux.
  • Experience with observability, distributed tracing, structured logging, metrics pipelines, and alerting platforms such as Cloud Monitoring, Datadog, and Prometheus/Grafana.
  • Experience defining and operating SLOs, SLIs, and error budgets in production.
  • Strong understanding of IAM, encryption, secrets management, network policies, compliance frameworks, multi-cloud networking, identity federation, and cost governance.
  • Fluency in at least one systems or scripting language: Python, Go, or Bash.
  • Experience with Kubernetes/GKE, service mesh, traffic management, data engineering patterns, data governance, data lineage, and platform-level data quality enforcement.
  • Understanding of agentic AI architectures and reliability challenges for LLM-based, event-driven, and multi-agent systems.
  • Experience leading without authority and communicating complex reliability concepts to technical and executive stakeholders.
  • Retail or e-commerce data platform experience, FinOps exposure, Azure-native services, cross-cloud identity management, or prior Staff/Principal-level SRE experience are desirable.

Benefits

  • Spain-based remote work arrangement.
  • Full-time position.
  • Opportunity to shape enterprise data and AI platform reliability for a global retail and supply chain organization.
  • Mentorship, cross-functional leadership, and organization-wide technical influence.

Tech Stack

AzureBashDatadogGoGoogle BigQueryGoogle Cloud PlatformGrafanaHelmKubernetesPrometheusPythonTerraform

Categories

DevOpsSite Reliability
Levi Strauss & Co

About Levi Strauss & Co

10,000+ employees

Levi Strauss & Co. designs, markets, and sells Levi’s denim and casual apparel, Dockers khakis, Signature/Denizen value lines, and Beyond Yoga activewear to consumers worldwide through company-owned stores, e-commerce, and wholesale retail partners. Founded in 1853 and headquartered in San Francisco, the company is publicly traded on the NYSE under the ticker LEVI. It is widely credited with inventing the blue jean in 1873.

Contact me