3 days ago
Hyderābād, IndiaSenior
Responsibilities
- Implement, manage, and maintain scalable, reliable infrastructure using infrastructure-as-code tools.
- Develop observability solutions to support service availability and performance.
- Design, build, and maintain CI/CD pipelines for deployment.
- Collaborate with development teams to improve service operability and reliability.
- Participate in a global on-call rotation for critical production systems.
- Automate processes and reduce operational toil.
- Manage and optimize cloud resources for cost efficiency and performance.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.
- 5+ years of experience as a Site Reliability Engineer, DevOps Engineer, or similar.
- Expertise in one or more programming languages such as Python or Go.
- Experience with cloud platforms such as AWS or GCP.
- Proficiency with configuration management and infrastructure-as-code tools such as Terraform or Salt.
- Proficiency with Kubernetes-based environments.
- Experience with monitoring and logging tools such as Prometheus, Grafana, or OpenSearch.
- Excellent technical communication skills.
- Preferred experience with large-scale distributed systems and knowledge of networking and security best practices.
- Preferred familiarity with SQL and NoSQL database technologies.
Benefits
- Work-from-office/in-office position based in Hyderabad, India.
- Participation in a global on-call rotation.
Tech Stack
Categories
Site Reliability
