1 hour ago
Bogotá, ColombiaMid Level
Responsibilities
- Create, respond to, and continuously improve platform alerts and runbooks.
- Monitor the platform, identify issues, and triage incidents.
- Perform impact assessments, communicate incident summaries, escalate incidents to owner teams, and act as Incident Leader when required.
- Execute mitigation actions to minimize impact and restore service.
- Write, test, secure, and maintain documented scripts and automation.
- Troubleshoot complex technical problems, collaborate with senior engineers, and ensure stable and reliable service delivery.
- Provide and receive constructive feedback to support continuous improvement.
Requirements
- Strong understanding of monitoring, alerting, and incident management processes.
- Hands-on scripting and automation experience.
- Solid troubleshooting and root cause analysis skills.
- Ability to work independently on complex technical topics.
- Clear and structured communication skills during incidents.
- Proactive reliability-focused mindset and comfort collaborating with cross-functional teams.
- Fluent spoken and written English.
Benefits
- Flexible hybrid work arrangements with remote work and flexible working hours.
- Employee Stock Ownership Plan with stock options.
- Time off, special leave days, and work-life balance and well-being support.
- Internal mobility, upskilling, mentorship, onboarding, and training programs.
- Short- and long-term international mobility opportunities in Infobip hubs.
- Compensation, performance-driven bonuses, and regular reviews aligned with experience, industry, and market standards.
Categories
Site Reliability
