8 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Architect scalable, secure, and cost-efficient cloud and hybrid infrastructure solutions and establish reusable engineering standards.
- Define organisation-wide monitoring, logging, alerting, SLI/SLO, error-budget, and observability practices.
- Set standards for high availability, disaster recovery, backup, failure testing, and operational excellence.
- Partner with development, security, and platform teams on CI/CD and GitOps workflows.
- Maintain technical documentation, runbooks, architectural decision records, and operational standards.
- Identify reliability bottlenecks, lead complex root-cause analysis, automate operational work, and improve platforms.
- Troubleshoot complex production issues, lead major incident response, provide third-line support, and participate in on-call escalation.
- Mentor SREs and engineers, lead enablement across Windows, Linux, AWS, and Kubernetes, and influence company-wide technical standards.
- Shape the SRE roadmap and lead prioritisation of reliability and platform programmes.
Requirements
- 10+ years of experience in SRE, DevOps, Platform Engineering, or Cloud Engineering, including senior or principal-level work.
- Deep hands-on experience with on-premise infrastructure and large-scale AWS environments.
- Expert production Kubernetes experience, preferably with EKS, EKS Anywhere, or EKS Hybrid, including cluster design, scaling, and hardening.
- Experience migrating IIS/.NET and Linux applications to cloud and Kubernetes.
- Expert Terraform infrastructure-as-code skills, including module design and standards for large teams.
- Knowledge of GitOps with ArgoCD, automated deployments, and configuration management.
- Broad experience across Windows, Linux, networking, and distributed systems.
- Experience with Datadog, Prometheus, OpenTelemetry, and SLI/SLO/error-budget practices.
- Experience applying security best practices in cloud and hybrid environments and operating in 24/7 mission-critical environments.
- Strong communication, cross-team influence, mentoring, and technical standards leadership skills.
Benefits
- Safe home pickup and home drop.
- Regular bonus and pension.
- 24 days of annual leave plus additional paid wellbeing and development days.
- Life assurance, income protection, private healthcare, and wellbeing support.
- INR 3,000 per month communication allowance.
- Up to INR 16,000 per year in Crèche expenses for children under 3.
- Entain India benefits and practical arrangements are provided for the role; no remote or office work arrangement is stated.
Tech Stack
Categories
DevOpsSite Reliability
About Entain
Entain is a London-headquartered, LSE-listed FTSE 100 sports betting and gaming group. It operates online sportsbooks, casino, poker and bingo, and runs retail betting shops, under brands including Ladbrokes, Coral, bwin, and a U.S. joint venture, BetMGM, with MGM Resorts. Revenue comes from consumer wagering across regulated markets, supported by proprietary technology platforms.