2 days ago
Remote, IndiaStaff+
Responsibilities
- Act as the L3 escalation point and technical SME for complex AWS, operating-system, Kubernetes, infrastructure, and deployment issues.
- Design and maintain infrastructure as code with Terraform and optionally CloudFormation, including reusable modules, state management, and security controls.
- Build and operate CI/CD pipelines and GitOps workflows with environment promotion, rollback, and drift remediation.
- Develop operational automation using Ansible/AWX, Bash, PowerShell, Python, and related tools.
- Design, deploy, and maintain Kubernetes workloads, Amazon EKS clusters, Helm charts, and environment-specific templates.
- Administer and secure Linux and Windows systems through hardening, patching, performance tuning, and automated deployments.
- Lead major incident response, root-cause analysis, disaster-recovery planning and testing, failover/failback validation, and permanent remediation.
- Define standards for infrastructure as code, Kubernetes, CI/CD, monitoring, logging, and operational processes.
- Improve platform observability, monitoring, alerting, logging, and operational reliability.
- Mentor L1 and L2 engineers, review and approve infrastructure changes, and provide hands-on technical guidance.
- Create and maintain SOPs, runbooks, architecture documentation, and disaster-recovery documentation.
- Collaborate with product and customer engineering teams on cloud solutions, migrations, and optimizations.
Requirements
- 10+ years of relevant experience in cloud engineering, infrastructure, DevOps, or a related field.
- Expert-level knowledge of AWS services, cloud architecture, security, networking, and operational best practices.
- Deep knowledge of Terraform module design, state management, Terragrunt patterns, and infrastructure-as-code security.
- Extensive Linux and Windows administration, automation, and security-hardening experience.
- In-depth Kubernetes architecture and operations experience, preferably with Amazon EKS, Helm, and GitOps patterns.
- Hands-on experience building CI/CD pipelines, managing artifacts, testing automation, and deployment automation.
- Experience with Ansible/AWX for configuration management, patching, application deployment, and operational tasks.
- Expert Linux administration and solid Windows Server administration skills, with Bash, PowerShell, and Python scripting.
- Strong Git skills, including branching strategies, code review, hooks, and access controls.
- Experience implementing monitoring, logging, and alerting with tools such as CloudWatch, Prometheus, and ELK/EFK.
- Knowledge of secure IAM practices, secret management, network security, encryption, and common frameworks including NIST, HIPAA, and PCI.
- Experience planning, testing, and documenting disaster recovery and resilient failover processes.
- Preferred AWS Solutions Architect or AWS DevOps Engineer Professional certification.
- CNCF/Kubernetes certifications such as CKA, CKAD, or CKS, and Azure or GCP certifications, are beneficial.
Tech Stack
AnsibleAWSAzureBashGitGitHub ActionsGitLab CI/CDGoogle Cloud PlatformHelmJenkinsKubernetesLinuxPowerShellPrometheusPythonTerraformWindows
Categories
About Rackspace
Rackspace provides managed cloud and IT services for enterprises, including multicloud operations, cloud migration, private cloud, managed security, and application/platform management across AWS, Azure, Google Cloud, and VMware. It operates a services-led business model with consulting and ongoing management for regulated and mission‑critical workloads, including healthcare. Founded in 1998 and headquartered in San Antonio, Texas, Rackspace is majority‑owned by Apollo Global Management.
