Rackspace

SRE Engineer II/III

Rackspace
Apply
20 hours ago
Hyderābād, IndiaMid Level / Senior

Responsibilities

  • Own L2/L3 incident response, root-cause analysis, and post-mortems for production issues.
  • Monitor system health and maintain SLA/SLO adherence.
  • Automate operational tasks and reduce repetitive toil through scripting and tooling.
  • Collaborate with development teams on deployment reliability and capacity planning.
  • Participate in an on-call rotation and maintain operational runbooks.
  • Build dashboards, configure meaningful alerts, and trace issues end-to-end.
  • Diagnose operating-system, infrastructure, network, HTTP, and application-level performance issues.

Requirements

  • 3–8 years of experience in site reliability, production operations, or a related engineering function.
  • Hands-on experience with Linux, including system and process management, file systems, and performance tuning.
  • Working knowledge of Windows Server and event log analysis.
  • Practical experience with AWS, Azure, or GCP compute, storage, IAM, networking, and managed services.
  • Familiarity with Terraform or equivalent infrastructure-as-code tools.
  • Proficiency in Python and Bash for automation, API interaction, and operational tooling.
  • Strong TCP/IP, DNS, TLS/SSL, NAT, firewall, routing, latency, packet-loss, and port-exhaustion troubleshooting skills.
  • Experience troubleshooting HTTP/HTTPS methods, status codes, headers, request lifecycles, access logs, and reverse proxies.
  • Experience with Prometheus, Grafana, Datadog, or ELK and with building dashboards and alerts.
  • Exposure to Ansible or similar configuration-management tools is preferred.
  • Kubernetes and Docker experience is preferred.
  • Familiarity with Kafka or RabbitMQ is preferred.
  • Basic troubleshooting experience with MySQL, PostgreSQL, or Redis is preferred.
  • Knowledge of ITIL fundamentals and ITSM tools such as Jira Service Management or ServiceNow is preferred.
  • Strong analytical thinking and clear communication under pressure are required.

Benefits

  • Hyderabad work-from-office position with a 24/7 work shift and on-call rotation.
  • Equal employment opportunity and accommodation support are provided.
  • Opportunity to work at Rackspace Technology, a multicloud solutions company recognized by Fortune, Forbes, and Glassdoor.

Tech Stack

AnsibleApache KafkaAWSAzureBashDatadogDockerGoogle Cloud PlatformGrafanaKubernetesLinuxMySQLPostgreSQLPostmanPrometheusPythonRabbitMQRedisTerraform

Categories

Site Reliability
Rackspace

About Rackspace

5,001-10,000 employees

Rackspace provides managed cloud and IT services for enterprises, including multicloud operations, cloud migration, private cloud, managed security, and application/platform management across AWS, Azure, Google Cloud, and VMware. It operates a services-led business model with consulting and ongoing management for regulated and mission‑critical workloads, including healthcare. Founded in 1998 and headquartered in San Antonio, Texas, Rackspace is majority‑owned by Apollo Global Management.

Contact me