1 month ago
Remote, ColombiaSenior
Responsibilities
- Operate, troubleshoot, and upgrade Kubernetes workloads and clusters in production.
- Design and maintain Azure Monitor, Log Analytics, KQL, Prometheus, and Grafana observability solutions.
- Define and implement service level indicators, service level objectives, and error budgets.
- Design alerting and incident response processes, author runbooks, and participate in on-call operations.
- Design and test backup, restore, and disaster recovery processes against RPO and RTO targets.
- Read and modify infrastructure as code using Terraform or Bicep and maintain Azure DevOps pipelines.
- Use Python, PowerShell, or Bash for scripting and automation.
Requirements
- 6+ years of experience, including 3+ years operating Kubernetes in production.
- Experience with site reliability engineering or production operations for Kubernetes workloads at scale.
- Experience with Azure Monitor, Log Analytics, KQL, workspace design, data collection rules, and retention strategy.
- Experience with Prometheus and Grafana metrics, exporters, recording rules, alerting rules, and dashboard design.
- Experience designing alerting, incident response, runbooks, on-call practices, backup and restore, and disaster recovery.
- Ability to read and modify Terraform or Bicep and Azure DevOps pipelines.
- Professional working English.
- Preferred experience with Azure Managed Prometheus, Azure Managed Grafana, OpenTelemetry, distributed tracing, Azure Backup, Azure Site Recovery, virtual machine snapshots, chaos engineering, Oracle observability, AKS cost and capacity management, incident tooling, and postmortems.
- Preferred certification such as CKA, AZ-400, or equivalent.
Benefits
- Remote work from Colombia.
- 13 floating holidays.
- 15 vacation days per completed year.
- Good working environment.
Tech Stack
Categories
Site Reliability
