1 day ago
Warsaw, PolandStaff+
Responsibilities
- Lead the development of SRE solutions for monitoring, alerting, anomaly detection, self-healing, and reliability testing.
- Design and implement reliability, scalability, performance, capacity planning, resource management, and automation strategies.
- Build tools that reduce manual operational work, improve developer experience, and increase system reliability.
- Define and manage SLOs and error budgets with engineering teams.
- Lead incident response, root-cause analysis, problem management, change management, disaster recovery, and platform readiness activities.
- Improve observability, operational excellence, application performance, infrastructure, onboarding pipelines, and production operations.
- Collaborate with product teams, clients, vendors, and stakeholders on architecture and operational initiatives.
- Mentor and guide junior SREs and promote knowledge sharing and continuous improvement.
- Participate in rotational on-call support, including regular and evening shifts and weekends as needed.
Requirements
- Bachelor’s or Master’s degree in Computer Science or a related field.
- At least 5 years of experience in SRE or cloud infrastructure leadership, with the posting describing a typical range of 5–8+ years.
- Extensive production-grade expertise in Microsoft Azure and cloud-native reliability practices.
- Expertise with Infrastructure as Code using Bicep, ARM, and Terraform.
- Experience with monitoring, logging, observability, distributed tracing, synthetic monitoring, and incident frameworks.
- Experience leading incident response and platform-wide reliability improvements using ITIL practices.
- Experience with identity management and secure access protocols including SAML, OAuth, and OIDC.
- Experience with networking, Kubernetes, Docker, APIs, scripting, and relational and cloud databases.
- Established ability to mentor engineers, guide architectural decisions, and connect long-term strategies with implementation.
- Knowledge of AI/ML-based anomaly detection and log aggregation and analysis tools.
- Experience managing onboarding projects and live production operations.
- Familiarity with SimCorp Dimension and Salesforce is a plus.
- Collaborative mindset, continuous-learning orientation, and ability to work across functions.
Benefits
- Hybrid work model with two days in the office and three days remote, plus flexible working hours.
- Rotational regular and evening shifts with weekend or on-call support as needed.
- Occasional remote work from Poland and internationally, subject to policy, with up to 24 domestic and 20 international days per year.
- Modern office near Wilanowska metro station with quiet zones and ergonomic workstations.
- Annual bonus structure and holiday allowance upon a two-week vacation.
- Employer-paid Medicover Platinum healthcare package, with employee-paid family upgrades available.
- Multisport card with 75% employer contribution.
- Unum group life insurance and Medicover travel insurance, with optional employee-paid upgrades where stated.
- Possibility to join the Deutsche Börse Group Share Plan after one year.
- Professional training, courses, language classes, career development opportunities, integration events, volunteering initiatives, and employee-led clubs.
Tech Stack
Categories
Site Reliability
About SimCorp
SimCorp builds SimCorp Dimension, a front-to-back investment management platform and managed services used by buy-side institutions such as asset managers, pension funds, and insurers. It sells software licenses and cloud-delivered operations (SaaS/managed services) covering portfolio management, trading, risk, accounting, and reporting. Founded in 1971 and headquartered in Copenhagen, it is a subsidiary of Deutsche Börse Group and reports serving 40 of the world’s top 100 financial companies.
