Senior Engineer - Site Reliability Engineering
MontyCloud20 days ago
Bengaluru, India or Hyderābād, IndiaSenior
Responsibilities
- Design and implement automation solutions for managing and monitoring cloud infrastructure and applications.
- Collaborate with SREs and cross-functional teams to proactively address reliability issues.
- Monitor cloud infrastructure and application health and performance, troubleshooting and optimizing systems.
- Participate in on-call rotations and respond to and resolve incidents.
- Lead and promote disaster recovery, chaos engineering, incident response, and reliability best practices.
- Conduct post-mortem analyses and incorporate lessons learned into future operations.
Requirements
- At least 5 years of experience as an SRE in a SaaS platform environment.
- At least 3 years of experience managing and optimizing SaaS platforms.
- At least 3 years of hands-on AWS experience.
- At least 4 years of experience with Ansible, Puppet, or Chef.
- At least 4 years of experience with Python or similar scripting languages.
- At least 3 years of experience with Splunk, New Relic, Datadog, AWS CloudWatch, or AWS X-Ray.
- At least 3 years of experience leading disaster recovery efforts and implementing chaos engineering practices.
- At least 4 years of on-call and incident management experience.
- At least 4 years of end-to-end application development experience.
- At least 3 years of experience leading post-mortem analysis sessions.
- Bachelor’s or master’s degree in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience.
- Strong problem-solving, communication, independent-working, and collaboration skills.