MontyCloud

Senior Engineer - Site Reliability Engineering

MontyCloud
Apply
20 days ago
Bengaluru, India or Hyderābād, IndiaSenior

Responsibilities

  • Design and implement automation solutions for managing and monitoring cloud infrastructure and applications.
  • Collaborate with SREs and cross-functional teams to proactively address reliability issues.
  • Monitor cloud infrastructure and application health and performance, troubleshooting and optimizing systems.
  • Participate in on-call rotations and respond to and resolve incidents.
  • Lead and promote disaster recovery, chaos engineering, incident response, and reliability best practices.
  • Conduct post-mortem analyses and incorporate lessons learned into future operations.

Requirements

  • At least 5 years of experience as an SRE in a SaaS platform environment.
  • At least 3 years of experience managing and optimizing SaaS platforms.
  • At least 3 years of hands-on AWS experience.
  • At least 4 years of experience with Ansible, Puppet, or Chef.
  • At least 4 years of experience with Python or similar scripting languages.
  • At least 3 years of experience with Splunk, New Relic, Datadog, AWS CloudWatch, or AWS X-Ray.
  • At least 3 years of experience leading disaster recovery efforts and implementing chaos engineering practices.
  • At least 4 years of on-call and incident management experience.
  • At least 4 years of end-to-end application development experience.
  • At least 3 years of experience leading post-mortem analysis sessions.
  • Bachelor’s or master’s degree in Computer Science, Engineering, or a related technical field, or equivalent hands-on experience.
  • Strong problem-solving, communication, independent-working, and collaboration skills.

Tech Stack

AnsibleAWSChefDatadogGitLab CI/CDJenkinsPuppetPythonSplunk

Categories

Site Reliability
Contact me