2 hours ago
Bucharest, RomaniaStaff+
Responsibilities
- Guide tribes in implementing Site Reliability Engineering practices and capabilities.
- Review and analyze service implementations for monitoring, alerting, and resilience architecture.
- Provide knowledge sessions on SLI/SLO definition, resilience testing, disaster recovery plans, and toil identification.
- Support root-cause analyses and post-mortems and identify actions to prevent recurring incidents.
- Analyze incident, problem, and change data to advocate for structural improvements.
- Provide hands-on support to application teams implementing monitoring and alerting standards.
- Contribute to the E&R organization and participate in the Global SRE Guild.
- Mentor and coach other SRE Experts.
Requirements
- Good understanding of Linux, preferably RHEL, or Windows.
- Good understanding of SQL and familiarity with relational databases such as Oracle and MS SQL.
- Good understanding of CI/DC standards, preferably Azure DevOps.
- Extensive knowledge of IT tools for sharing and collaboration.
- Experience with monolithic application landscapes.
- User experience with monitoring and alerting tools such as Prometheus, ELK, and Grafana.
- Familiarity with SRE concepts including SLI, SLO, error budgets, and availability reporting.
- Understanding of consumers and engineers, strong problem-solving ability, a strong voice, and healthy skepticism.
- Good knowledge of IT security principles, containers, and cloud tooling.
- Knowledge of OpenTelemetry is a bonus.
Benefits
- Flexible working arrangements.
- Role based in or connected to ING Hubs Romania locations in Bucharest and Cluj-Napoca.
- Collaborative environment with distributed teams and participation in global SRE guilds.
