2 hours ago
Bucharest, RomaniaSenior
Responsibilities
- Review monitoring, observability, alerting, resilience architecture, and operational readiness for services.
- Implement and improve monitoring, alerting, logging, and observability solutions.
- Define, implement, and monitor SLIs, SLOs, error budgets, and availability reporting.
- Support root cause analyses, post-mortems, incident investigations, and structural reliability improvements.
- Troubleshoot complex production issues and help engineering teams improve service stability.
- Identify operational toil and contribute to automation and operational excellence initiatives.
- Support resilience testing, disaster recovery exercises, and operational readiness assessments.
- Collaborate with development, architecture, platform, and operations teams throughout the software lifecycle.
- Improve runbooks, operational processes, and engineering practices, and share knowledge through documentation and workshops.
- Contribute to the E&R organization, global SRE guilds, and communities of practice.
Requirements
- Extensive knowledge of Linux, preferably RHEL, or Windows.
- Good knowledge of SQL and familiarity with Oracle or Microsoft SQL databases, or NoSQL databases such as Cassandra.
- Good knowledge of CI/CD standards, preferably Azure DevOps.
- Experience with monolithic and distributed application landscapes.
- Good knowledge of at least one programming language.
- Good knowledge of networking and IPv4 and IPv6 stacks.
- Experience with monitoring, observability, and alerting tools such as Prometheus, ELK, Grafana, and OpenTelemetry.
- Familiarity with SRE concepts including SLI, SLO, error budgets, reliability engineering, and availability reporting.
- Good knowledge of containers and cloud tooling.
- Knowledge and experience using OpenTelemetry is a bonus.
- Good knowledge of IT security principles.
- Ability to translate customer and engineering needs into practical solutions and collaborate within cross-functional teams.
- Advanced English proficiency and a continuous-improvement mindset.
Benefits
- Flexible and highly collaborative working environment.
- Opportunity to support payment and settlement services used across multiple ING business lines and participate in global SRE communities.
Categories
Site Reliability
