2 days ago
Remote, MexicoStaff+
Responsibilities
- Own the architecture and roadmap for a multi-cluster Kubernetes platform, including scaling, upgrades, and multi-tenancy.
- Establish GitOps-based deployment workflows using Argo CD or Flux.
- Design and evolve an internal developer platform with self-service access to compute, environments, and observability.
- Architect service mesh, networking, and ingress strategies for reliable and secure service communication.
- Define standards for GPU and ML workload scheduling and resource management on Kubernetes.
- Partner with AI/ML teams on infrastructure requirements for training and inference workloads.
- Drive capacity planning, resource optimization, and cost optimization across Kubernetes and cloud infrastructure.
- Set technical direction and review architecture for platform-impacting engineering changes.
- Own platform reliability through SLOs, incident-response leadership, and postmortems.
- Mentor experienced engineers and represent platform engineering in cross-organizational technical decisions.
- Influence teams toward common platform standards without direct reporting authority.
Requirements
- Bachelor’s degree in Computer Science or a related field, or equivalent work experience.
- 8+ years of DevOps, platform, or infrastructure engineering experience.
- 5+ years of hands-on production Kubernetes experience, including cluster architecture, upgrades, and multi-tenant environments.
- Strong hands-on experience operating cloud infrastructure across AWS and Google Cloud Platform.
- Strong GitOps experience with Argo CD or Flux.
- Strong Infrastructure-as-Code experience with Terraform or Pulumi.
- Deep understanding of container orchestration, networking, ingress, and service mesh technologies such as Istio, Linkerd, or Cilium.
- Experience operating Nginx, Kafka, and Redis at scale.
- Experience with GPU scheduling, node pools, and resource quotas for ML/AI workloads in GKE and EKS.
- Experience designing internal developer platforms or paved-road tooling for engineering organizations.
- Strong Linux, networking, storage, and security fundamentals.
- Experience operating observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry at platform scale.
- Ability to communicate effectively and influence technical direction across teams.
- Experience calculating, forecasting, and optimizing cloud and Kubernetes infrastructure costs and evaluating architecture, capacity, compute, and resource-allocation decisions.
- Experience with capacity planning and resource optimization while balancing performance, reliability, scalability, and cost.
- Additional helpful experience includes Kubernetes policy-as-code and security tooling, image scanning, and AWS, GCP, or Azure certifications, including CKA or CKS.
- Experience designing infrastructure for ML/AI training or inference workloads and understanding GPU scheduling, resource allocation, node pools, quotas, reliability, and cost considerations.
Benefits
- Remote work is available across Mexico; employees within 80 kilometers of the Guadalajara office are expected to work onsite.
- Health insurance, including vision and dental coverage, for employees and eligible family members.
- Life insurance equal to 24 times the monthly salary.
- New full-time employees receive 10 personal days before their first anniversary, 12 vacation days on their first anniversary, and 5 personal days annually thereafter in addition to vacation time.
- 30-day Christmas bonus, 50% vacation premium, company-matched food vouchers, and a 13% matched savings fund.
- Employee Assistance Program and wellness initiatives.
- Ongoing learning, development, and career advancement opportunities.
Tech Stack
Apache KafkaArgo CDAWSAzureDatadogGoogle Cloud PlatformGrafanaIstioKubernetesLinuxPrometheusRedisTerraform
