20 hours ago
Hyderābād, IndiaMid Level / Senior
Responsibilities
- Own L2 incident response, root cause analysis, and post-mortems for production issues.
- Monitor system health and maintain SLA/SLO adherence.
- Automate operational tasks and reduce repetitive toil using scripting and tooling.
- Collaborate with development teams on deployment reliability and capacity planning.
- Participate in the on-call rotation and maintain operational runbooks.
- Troubleshoot operating systems, networks, HTTP applications, observability data, and infrastructure-level bottlenecks.
- Build dashboards, configure alerts, and trace issues end to end.
Requirements
- 3–8 years of experience in site reliability engineering or related infrastructure operations.
- Hands-on Linux experience, including RHEL or Ubuntu, systemd, process management, file systems, and performance tuning.
- Working knowledge of Windows Server and event log analysis.
- Practical experience with AWS, Azure, or GCP compute, storage, IAM, networking, and managed services.
- Familiarity with Terraform or equivalent infrastructure-as-code tools.
- Proficiency in Python and Bash for automation, API interaction, and operational tooling.
- Strong TCP/IP, DNS, TLS/SSL, NAT, firewall, routing, packet-flow, and network troubleshooting knowledge.
- Experience debugging HTTP/HTTPS behavior using curl, Postman, access logs, and reverse proxy configurations.
- Experience with Prometheus, Grafana, Datadog, or ELK and the ability to create dashboards and meaningful alerts.
- Kubernetes, Docker, Kafka, RabbitMQ, MySQL, PostgreSQL, Redis, ITIL, Jira Service Management, or ServiceNow experience is a plus.
- Strong analytical thinking and clear communication under pressure.
Benefits
- Work from the office in Hyderabad.
- 24/7 work-shift coverage with participation in an on-call rotation.
- Equal employment opportunity and workplace accommodation support are provided.
Tech Stack
AnsibleApache KafkaAWSAzureBashDatadogDockerGoogle Cloud PlatformGrafanaKubernetesLinuxMySQLPostgreSQLPostmanPrometheusPythonRabbitMQRedisTerraform
Categories
Site Reliability
About Rackspace
Rackspace provides managed cloud and IT services for enterprises, including multicloud operations, cloud migration, private cloud, managed security, and application/platform management across AWS, Azure, Google Cloud, and VMware. It operates a services-led business model with consulting and ongoing management for regulated and mission‑critical workloads, including healthcare. Founded in 1998 and headquartered in San Antonio, Texas, Rackspace is majority‑owned by Apollo Global Management.
