4 days ago
Remote, Argentina +2 moreSenior
Responsibilities
- Rearchitect and codify AWS and Azure infrastructure with OpenTofu/Terraform, Ansible, reusable modules, remote state, and GitOps-based pipelines.
- Audit cloud environments against AWS and Azure Well-Architected Frameworks and implement networking, IAM, landing zone, account structure, cost governance, and policy guardrails.
- Define FinOps standards, cost allocation dashboards, rightsizing recommendations, and reserved capacity planning across both clouds.
- Build the observability stack for metrics, logs, traces, dashboards, OpenTelemetry instrumentation, SLIs, SLOs, error budgets, and burn-rate alerting.
- Establish SRE practices including incident runbooks, post-incident reviews, chaos engineering exercises, service instrumentation, alerting, and capacity planning.
- Deploy and operate an Internal Developer Platform and developer portal with service catalogs, scaffolding templates, runbooks, API documentation, and on-call ownership.
- Create golden paths for service creation, Kubernetes deployment, database provisioning, secrets management, and CI/CD setup.
- Own CI/CD platform layers, reusable workflow libraries, container build and scan pipelines, environment promotion, and security scanning.
- Operate EKS and/or AKS clusters and establish standards for Helm, admission controllers, RBAC, network policies, autoscaling, and service mesh.
- Build self-service provisioning through Backstage scaffolder actions and Terraform automation.
- Support workload migrations, embed security into platform work, write architecture documentation and runbooks, and mentor junior platform engineers.
Requirements
- 5+ years of platform, infrastructure, or DevOps engineering experience with direct production ownership on AWS and/or Azure.
- Deep OpenTofu/Terraform experience covering module authoring, state management, workspace strategy, remote backends, and CI/CD integration.
- Strong Kubernetes operations experience with EKS and/or AKS, Helm, admission controllers, RBAC, network policies, and autoscaling.
- Hands-on experience with at least two observability technologies including Prometheus, Grafana, Loki, Tempo, Datadog, or OpenTelemetry, plus SLI/SLO and alert engineering.
- Experience authoring GitHub Actions pipelines, designing reusable workflows, and owning container build and scan pipelines.
- GitOps experience with ArgoCD or Flux and familiarity with progressive delivery patterns such as canary and blue-green deployments.
- Internal Developer Platform experience with Backstage or an equivalent developer portal, GitHub, scaffolding templates, service catalogs, or self-service provisioning tools.
- Experience with policy-as-code, IaC scanning, secrets management, container hardening, and shift-left security practices.
- Strong communication and documentation skills, including presenting architecture decisions to engineering peers and leadership.
- Desirable qualifications include Terramate, AWS FIS, Chaos Monkey, error budget management, incident command, capacity planning, Istio or Linkerd, FinOps tools such as Kubecost or CloudHealth, AI/ML infrastructure, Python or Go, and relevant AWS, Azure, Kubernetes, or Terraform certifications.
Benefits
- Company-paid medical insurance, mental health programs, and five undocumented sick-leave days per year.
- Internal events, meetups, conferences, workshops, Udemy access, language courses, and company-paid certifications.
- 20 working days of paid vacation and local bank holidays with long-term employment.
- 100% remote work mode.
- Internal Mobility Program and opportunities to change projects for professional growth.
- International projects, a global professional community, and regular team-building events.
Tech Stack
Categories
DevOpsSite Reliability
About Ciklum
Ciklum is an IT services firm that builds custom digital products and provides nearshore engineering teams, QA, DevOps, data analytics, and UX for enterprises and high-growth startups. It operates a services model spanning product engineering, e-commerce solutions, and managed delivery centers. Founded in 2002 and headquartered in London, the privately held company serves global clients including Just Eat and Zurich Insurance.
