IG Group

Senior Platform SRE

IG Group
Apply
25 days ago
Kraków, PolandSenior

Responsibilities

  • Build and own the reliability platform across IG’s AWS and HashiCorp Nomad estate.
  • Implement monitoring and observability using OpenTelemetry and distributed tracing, including instrumentation of Java and Python services.
  • Define and maintain SLOs, SLIs, error budgets, multi-window burn-rate alerts, and customer-journey reliability measures.
  • Establish operational readiness through automated deployments, blue/green and canary releases, automated rollback, zero-downtime patching, and DORA metrics.
  • Engineer self-healing capabilities including auto-remediation, error-budget-gated rollback, automated traffic rerouting, and fault-tolerance patterns.
  • Design and execute controlled chaos experiments across AWS using hypothesis-driven failure scenarios.
  • Build automation tools and CI/CD pipelines while contributing production-quality application code.
  • Contribute to IG’s SRE AI agent for incident investigation and reliability review.
  • Author and evolve SRE standards, including SLO methodology, error-budget policy, observability guidance, and Production Readiness Review checklists.
  • Mentor junior SREs, developers, and Reliability Champions on production engineering and reliability patterns.
  • Guide teams on system design, capacity planning, architectural reviews, and closing observability gaps.
  • Own incident response, facilitate blameless post-incident reviews, maintain the Lessons Register, and track remediation actions to completion.

Requirements

  • 6+ years of experience across the required observability, SRE, CI/CD, container orchestration, software engineering, distributed systems, incident management, and chaos engineering areas.
  • Hands-on OpenTelemetry experience covering spans, metrics, traces, and context propagation, plus production use of Honeycomb, Datadog, Dynatrace, or Grafana.
  • Proven experience designing meaningful SLIs, setting error budgets, configuring multi-window burn-rate alerts, and partnering with development teams on reliability measurement.
  • Experience building safe release pipelines with blue/green or canary releases, automated rollback, and DORA metrics integration.
  • Required Kubernetes experience with EKS, AKS, or GKE; HashiCorp Nomad experience is advantageous.
  • Solid understanding of cloud networking and infrastructure as code, with Terraform preferred.
  • Production-quality Java and/or Python coding experience and the ability to modify application codebases to implement reliability patterns.
  • Strong understanding of distributed-system failure modes and patterns including circuit breakers, bulkheads, idempotency, graceful degradation, and load shedding.
  • Production on-call experience, blameless post-incident review facilitation, contributing-factor analysis, and remediation tracking; PagerDuty and ServiceNow familiarity is helpful.
  • Experience designing and executing controlled chaos experiments using AWS FIS, Gremlin, or an equivalent tool.
  • Experience in high-throughput production environments such as financial services or trading platforms is preferred.
  • Demonstrated ability to improve reliability and performance at scale and collaborate with development teams on observability improvements.
  • Strong troubleshooting, systems thinking, communication, technical enablement, automation, and manual-toil reduction skills.
  • Comfort writing RFCs, presenting at engineering forums, and developing standards that engineering teams adopt.

Benefits

  • Hybrid working model with three days in the office.
  • Tailored development programs, mentoring opportunities with leaders, and clear career progression.
  • Committees, sports clubs, and social clubs to expand professional and social networks.
  • Extra time off for volunteering and community work.

Tech Stack

Categories

DevOpsSite Reliability
IG Group

About IG Group

1,001-5,000 employees
Contact me