24 days ago
Remote, WorldwideSenior
Responsibilities
- Design, implement, and maintain advanced monitoring solutions for infrastructure and applications.
- Lead complex incident response, root cause analysis, and post-incident reviews for critical reliability issues.
- Develop automation scripts and tools to improve reliability, operational efficiency, and scalability.
- Analyze performance data to identify trends, anomalies, and potential system issues.
- Maintain system documentation and contribute to team knowledge sharing.
- Lead reliability engineering projects and collaborate with cross-functional teams.
- Mentor junior and mid-level engineers and support their technical development.
- Implement security measures and ensure compliance with security policies and industry standards.
- Drive continuous improvement and contribute to strategic infrastructure and operations planning.
- Conduct capacity planning and develop disaster recovery plans for business continuity.
- Participate in change management processes for system updates and changes.
- Promote an autonomous and inclusive work culture aligned with Spin's values.
Requirements
- Bachelor's degree in computer science, information technology, or a related field, or equivalent work experience.
- At least 7 years of experience in site reliability engineering or related fields.
- Deep understanding of system reliability, advanced monitoring, automation, and incident response.
- Proficiency with multiple scripting languages and automation tools.
- Experience with cloud platforms and containerization technologies.
- Strong problem-solving, troubleshooting, communication, teamwork, strategic thinking, and decision-making skills.
- Proven leadership and mentorship abilities.
- Willingness to learn and adapt to new technologies and processes.
Categories
Site Reliability
