about 5 hours ago
Responsibilities
- Work with teams to define SLIs and SLOs.
- Create systems for observability.
- Analyze failure scenarios and propose mitigations.
- Assist in creating runbooks for failure remediation.
- Reduce non-value-adding work.
- Participate in incident management and on-call duty.
Requirements
- 5 years of experience in software engineering, DevOps, QA, or cloud engineering.
- At least 2 years as a dedicated Site Reliability Engineer.
- Strong leadership and decision-making skills.
- Experience with incident management in high-traffic production environments.
- Proficient in programming and scripting.
- Basic knowledge of serverless services in public cloud providers.
- Extensive experience with monitoring systems like Datadog and Grafana.
- Familiarity with pipelining tools such as GitHub and Jenkins.
- Knowledge of microservices technologies like Docker and Kubernetes.
- Good conceptual understanding of software architecture.
- Excellent command of English (C1 or above).
- Experience with eCommerce platforms and international teams.
Benefits
- 24 working days of paid vacation.
- National holidays covered.
- Sick leave (up to 20/year).
- Medical insurance.
- Multisport card or Multikafeteria.
- Support for maternity and paternity leave.
- Internal workshops and learning initiatives.
- Professional certifications reimbursement.
- Mentoring program for career growth.
