about 7 hours ago
Responsibilities
- Understand and monitor application performance to identify issues.
- Analyze current systems to reduce existing problems and suggest upgrades.
- Provide support in monitoring, processes, tools, architecture, and Root Cause Analysis.
- Develop and maintain monitoring and alerting systems.
- Automate routine tasks to enhance system efficiency.
- Troubleshoot and resolve incidents and outages.
- Create automation scripts and tools for system management.
- Identify areas for improvement and design scalable solutions.
- Monitor alerts to prevent production outages.
- Maintain Run books for alerts.
Requirements
- 6-9 years of sysadmin experience with large-scale distributed systems.
- Strong foundation in cloud management.
- Proficiency in Unix shells, Python, and Go programming.
- Experience with MySQL or PostgreSQL databases.
- Excellent interpersonal and communication skills.
- Strong debugging and troubleshooting abilities.
- Experience with observability tools like Prometheus and Grafana.
- Hands-on experience with AWS and AWS-CLI.
- Experience in orchestration and containerization (Kubernetes, Containers).
- Familiarity with CI/CD tools like Jenkins and ArgoCD.
- Solid understanding of networking concepts.
- Experience with Linux OS and shell/Python scripting.
- Knowledge of API Gateway systems like Kong and Nginx.
- Understanding of security best practices.
- BS degree in Computer Science or related field.