
Site Reliability Engineer
Levi Strauss & Co3 months ago
Mexico City, MexicoSenior
Responsibilities
- Monitor production systems using dashboards, alerts, logs, metrics, and tracing to detect and triage issues.
- Participate in on-call rotations, respond to incidents using runbooks, and contribute to blameless post-mortems.
- Maintain SLO dashboards and alerting thresholds and help improve production reliability.
- Build scripts, tooling, and CI/CD pipeline components to reduce operational toil and improve deployment reliability.
- Operate and maintain GCP workloads including GKE, Cloud Run, BigQuery, Pub/Sub, GCS, and Composer.
- Manage infrastructure changes using Terraform and Helm and support self-service infrastructure initiatives.
- Support multi-cloud standards across GCP and Azure.
- Apply IAM, secrets management, encryption, network policy, and audit logging practices.
- Collaborate with Data Engineering, AI Platform, and Software Engineering teams on reliability from design through deployment.
- Participate in reliability reviews, design discussions, team ceremonies, and technical knowledge-sharing.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- At least 6 years of production experience in Site Reliability Engineering, DevOps, or Platform/Infrastructure Engineering.
- Hands-on experience with GCP services, particularly GKE, Cloud Run, BigQuery, Pub/Sub, and GCS.
- Working proficiency with Infrastructure-as-Code tools such as Terraform or Helm.
- Familiarity with observability tooling and metrics, logging, tracing, and alerting.
- Understanding of SLO/SLI concepts, error budgets, toil tracking, and on-call operations.
- Exposure to data security fundamentals including IAM, encryption, secrets management, and network policies.
- Proficiency in at least one scripting or systems language: Python, Bash, or Go.
- Experience with Kubernetes or GKE and containerized workload operations.
- Basic understanding of CI/CD pipelines and GitOps workflows such as ArgoCD or GitHub Actions.
- Familiarity with batch or streaming data pipelines and multi-cloud concepts across GCP and Azure.
- Strong communication skills for documenting incidents, runbooks, and technical processes.
- Preferred experience includes retail, e-commerce, or consumer goods; AI/ML platform operations; model-serving infrastructure monitoring; and FinOps or cloud cost visibility tooling.
Benefits
- Full-time position located in Mexico City, Mexico.
- Team knowledge-sharing, documentation, and hands-on development opportunities are provided.
Tech Stack
AzureBashDatadogGitHub ActionsGoGoogle BigQueryGoogle CloudGoogle Cloud PlatformGrafanaHelmKubernetesPrometheusPythonTerraform
Categories
Site Reliability