1 month ago
Responsibilities
- Own and evolve AWS infrastructure using Terraform and Terragrunt.
- Architect and scale AWS environments and manage containerized workloads with Kubernetes and Docker.
- Contribute to high-availability and disaster-recovery architecture and platform strategy.
- Lead deployment and release processes using Argo, Bitbucket, and Jenkins.
- Define and enforce SLOs, SLIs, and error budgets while reducing platform toil.
- Drive monitoring, dashboards, and alerting through Datadog, with Prometheus and Grafana among the observability tools.
- Build self-service internal developer platforms that help engineering teams ship faster.
- Own infrastructure projects end-to-end, from defining success criteria through execution and outcome measurement.
- Partner with engineering teams on long-term technical planning, document work, provide peer cross-training, and resolve JIRA tickets across cloud, CI/CD, deployments, and monitoring.
- Introduce AI-assisted engineering practices, including Claude and MCP integrations, into daily workflows.
Requirements
- At least six years of DevOps or SRE engineering experience in a cloud environment.
- Hands-on production experience with AWS, Kubernetes, containerization, and Terraform or Terragrunt.
- Strong Bash scripting skills and a deep understanding of SRE principles, including SLOs, SLIs, error budgets, toil reduction, and blameless post-mortems.
- Strong incident management and on-call experience, with an understanding of APIs, microservices, and distributed systems.
- Experience leading projects end-to-end, including defining success criteria, delivery, and outcome measurement.
- Ability to communicate effectively across teams and drive long-term technical planning.
- Practical experience with AI-assisted engineering tools such as Claude or Cursor and MCP-style integrations is a strong plus.
- Experience building AI/ML infrastructure, including model deployment and inference pipelines, is a plus.
Tech Stack
Categories
DevOpsSite Reliability
