2 days ago
Remote, IndiaSenior
Responsibilities
- Act as the L3 escalation point for complex AWS cloud, DevOps, Kubernetes, infrastructure, and deployment issues.
- Lead major technical activities, complex infrastructure changes, upgrades, migrations, platform improvements, and production deployments.
- Design and maintain infrastructure using Terraform and reusable Terraform modules.
- Develop reusable Ansible roles, playbooks, and AWX automation workflows for configuration, patching, deployments, and infrastructure operations.
- Design and troubleshoot ArgoCD-based GitOps deployments, including synchronization, drift, rollback, and environment promotion issues.
- Design and maintain Kubernetes workloads on Amazon EKS and reusable Helm charts.
- Improve Jenkins and other deployment automation processes and define deployment and operational standards.
- Lead disaster recovery planning, restore testing, failover/failback, recovery validation, and runbook development.
- Lead complex production incidents, perform root cause analysis, and implement permanent automated solutions.
- Improve monitoring, logging, alerting, and operational reliability.
- Define technical standards, reusable templates, modules, pipelines, and automation frameworks.
- Perform technical reviews of L1/L2 work and mentor engineers.
- Create and maintain SOPs, runbooks, technical standards, architecture documentation, disaster recovery documentation, and operational procedures.
Requirements
- 8–12 years of relevant experience in cloud engineering, infrastructure, DevOps, or a related field.
- Strong hands-on expertise with AWS, Amazon EKS, Kubernetes, Terraform, Terragrunt, AWX, Ansible, ArgoCD, Helm, Docker, Jenkins, Git, Linux, Windows, Bash, and Python.
- Experience with infrastructure automation, GitOps, CI/CD deployment processes, disaster recovery, monitoring and logging, production troubleshooting, and root cause analysis.
- Ability to independently lead technical activities, own complex cloud and DevOps changes, define standards, build automation, and lead disaster recovery activities.
- Ability to develop and review Terraform code, implement ArgoCD-based GitOps, automate with AWX/Ansible, handle complex production incidents, and mentor L1/L2 engineers.
About Rackspace
Rackspace provides managed cloud and IT services for enterprises, including multicloud operations, cloud migration, private cloud, managed security, and application/platform management across AWS, Azure, Google Cloud, and VMware. It operates a services-led business model with consulting and ongoing management for regulated and mission‑critical workloads, including healthcare. Founded in 1998 and headquartered in San Antonio, Texas, Rackspace is majority‑owned by Apollo Global Management.
