25 days ago
Kraków, PolandSenior
Responsibilities
- Build and own the reliability platform across IG’s AWS and HashiCorp Nomad estate.
- Implement monitoring and observability using OpenTelemetry and distributed tracing, including instrumentation of Java and Python services.
- Define and maintain SLOs, SLIs, error budgets, multi-window burn-rate alerts, and customer-journey reliability measures.
- Establish operational readiness through automated deployments, blue/green and canary releases, automated rollback, zero-downtime patching, and DORA metrics.
- Engineer self-healing capabilities including auto-remediation, error-budget-gated rollback, automated traffic rerouting, and fault-tolerance patterns.
- Design and execute controlled chaos experiments across AWS using hypothesis-driven failure scenarios.
- Build automation tools and CI/CD pipelines while contributing production-quality application code.
- Contribute to IG’s SRE AI agent for incident investigation and reliability review.
- Author and evolve SRE standards, including SLO methodology, error-budget policy, observability guidance, and Production Readiness Review checklists.
- Mentor junior SREs, developers, and Reliability Champions on production engineering and reliability patterns.
- Guide teams on system design, capacity planning, architectural reviews, and closing observability gaps.
- Own incident response, facilitate blameless post-incident reviews, maintain the Lessons Register, and track remediation actions to completion.
Requirements
- 6+ years of experience across the required observability, SRE, CI/CD, container orchestration, software engineering, distributed systems, incident management, and chaos engineering areas.
- Hands-on OpenTelemetry experience covering spans, metrics, traces, and context propagation, plus production use of Honeycomb, Datadog, Dynatrace, or Grafana.
- Proven experience designing meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and partnering with development teams on reliability measurement.
- Experience building safe release pipelines with blue/green or canary releases, automated rollback, and DORA metrics integration.
- Required Kubernetes experience with EKS, AKS, or GKE; HashiCorp Nomad experience is advantageous.
- Solid understanding of cloud networking and infrastructure as code, with Terraform preferred.
- Production-quality Java and/or Python coding experience and the ability to modify application codebases to implement reliability patterns.
- Strong understanding of distributed-system failure modes and patterns including circuit breakers, bulkheads, idempotency, graceful degradation, and load shedding.
- Production on-call experience, blameless post-incident review facilitation, contributing-factor analysis, and remediation tracking; PagerDuty and ServiceNow familiarity is helpful.
- Experience designing and executing controlled chaos experiments using AWS FIS, Gremlin, or an equivalent tool.
- Experience in high-throughput production environments such as financial services or trading platforms is preferred.
- Demonstrated ability to improve reliability and performance at scale and collaborate with development teams on observability improvements.
- Strong troubleshooting, systems thinking, communication, technical enablement, automation, and manual-toil reduction skills.
- Comfort writing RFCs, presenting at engineering forums, and developing standards that engineering teams adopt.
Benefits
- Hybrid working model with three days in the office.
- Tailored development programs, mentoring opportunities with leaders, and clear career progression.
- Committees, sports clubs, and social clubs to expand professional and social networks.
- Extra time off for volunteering and community work.