
Associate Senior Site Reliability Engineer
Total System Services, LLC (TSYS)1 hour ago
Pune, IndiaSenior
Responsibilities
- Improve availability, latency, performance, efficiency, capacity planning, monitoring, and emergency response for production systems.
- Design and execute chaos-engineering tests, game days, and performance experiments, then implement remediation plans.
- Build reliable, resilient, self-monitoring, and self-healing systems and improve automation and self-service through DevOps and GitOps practices.
- Develop monitoring and alerting systems for production and non-production virtualized infrastructure.
- Review designs for platform stability and risk, troubleshoot system and network issues, and support the Technical Operations Team.
- Evolve SDLC practices and tooling to incorporate site reliability considerations.
- Develop runbooks and improve technical documentation.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Business, Management Information Systems, or a related field.
- Typically at least 2 years of relevant experience.
- Experience with public and private clouds, Jenkins, Terraform, Ansible, OpenShift, Kubernetes, or AWS EKS.
- Ability to solve moderately scoped problems, exercise judgment within defined procedures, and build productive internal and external working relationships.
Tech Stack
Categories
Site Reliability