1 day ago
Remote, India or Bengaluru, IndiaSenior
Responsibilities
- Solve complex reliability and infrastructure problems through proactive troubleshooting, automation, and systems programming.
- Deploy and maintain observability platforms and internal tooling.
- Partner across teams to improve the reliability, scalability, and usability of products and services.
- Guide engineers and developers in evaluating and improving service performance.
- Investigate and troubleshoot complex issues with support, operations, and engineering teams.
Requirements
- Bachelor’s degree in Computer Science or Engineering.
- Six years of experience in Site Reliability Engineering or a related engineering role.
- Expertise in Linux system administration and understanding of TCP/IP, DNS, routing and switching, and storage concepts.
- Hands-on experience operating and troubleshooting containerized production systems with Kubernetes.
- Experience with Jenkins, Git, Prometheus, and Grafana.
- Experience with Infrastructure as Code and configuration management using Terraform, Ansible, SaltStack, or similar tools.
- Automation and scripting skills using Python, Bash, Go, Rust, or similar languages.
- Exposure to cloud storage systems and knowledge of CI/CD and DevOps practices.
Benefits
- Health, well-being, financial, and other employee benefits are provided.
- Akamai’s FlexBase program supports working from home, in an office, or in a combination of both.
Categories
Site Reliability
About Akamai
Akamai builds content delivery, cloud security, and edge compute services for enterprises that run consumer and business web, API, and video workloads. The public company (NASDAQ: AKAM), founded in 1998 and headquartered in Cambridge, MA, operates a global CDN and security platform and sells via subscriptions and usage-based services. It expanded into cloud infrastructure by acquiring Linode in 2022, offering IaaS and managed Kubernetes.
