about 5 hours ago
Responsibilities
- Work with teams to define SLIs and SLOs.
- Create systems for observability.
- Analyze failure scenarios and propose mitigations.
- Assist in creating runbooks for failure remediation.
- Reduce non-value-adding work.
- Participate in incident management and on-call duty.
Requirements
- 5 years of experience in software engineering, DevOps, QA, or cloud engineering.
- At least 2 years as a dedicated Site Reliability Engineer.
- Strong leadership and decision-making skills.
- Experience with incident management in high-traffic production environments.
- Proficiency in programming and scripting.
- Basic knowledge of serverless services in public cloud providers.
- Extensive experience with monitoring systems like Datadog and Grafana.
- Familiarity with pipelining tools such as GitHub and Jenkins.
- Knowledge of microservices technologies like Docker and Kubernetes.
- Excellent command of English (C1 or above).
- Experience with eCommerce platforms and international teams.
Benefits
- Flexibility with remote and hybrid work options.
- Career advancement opportunities with international mobility.
- Access to cutting-edge tools and professional development programs.
