
Senior Site Reliability Engineer
AutoRABIT Holding Inc.2 months ago
Hyderābād, IndiaSenior
Responsibilities
- Define and manage SLIs, SLOs, SLAs, and error budgets.
- Own incident response, on-call operations, root-cause analysis, and postmortems.
- Build and maintain observability for logs, metrics, alerts, and dashboards using Elasticsearch and Kibana.
- Automate infrastructure with Terraform and operate workloads across AWS ECS, EKS, EC2, Lambda, and SSM.
- Manage CI/CD pipelines using CodeBuild, CodeDeploy, and CodePipeline.
- Implement CIS Level 2-aligned security controls, OS hardening, and Trend Micro endpoint/workload security.
- Drive cost, performance, capacity, self-healing, auto-remediation, and runbook automation improvements.
- Support DynamoDB, PostgreSQL, and MySQL operations and performance tuning.
- Leverage Amazon Bedrock for automation and operational-intelligence use cases.
- Adhere to established internal controls.
Requirements
- Bachelor’s degree in computer science or a related field.
- 6+ years of experience in SRE, DevOps, or platform engineering.
- 3+ years of experience in AWS-based production environments.
- Strong experience with Kubernetes, including EKS, and Docker.
- Deep AWS experience with ECS, EKS, EC2, Lambda, IAM, VPC, and SSM.
- Hands-on experience with Terraform and CI/CD pipelines.
- Expertise with the ELK stack for log ingestion, parsing, alerting, and dashboards.
- Experience with Ubuntu and Amazon Linux administration and hardening.
- Knowledge of DNS, TCP/IP, and load balancing.
- Scripting experience with Python and Bash.
- Database operations and performance-tuning experience.
- Experience with CIS benchmarks and SOC 2/ISO controls.
- Experience designing multi-region and highly available systems.
- Experience with AIOps or AI-assisted observability.
Benefits
- Hybrid work arrangement in Hyderabad with 3 days from the office.
- Compensation of 24–28 LPA.
Tech Stack
Categories
Site Reliability