1GLOBAL

Senior Site Reliability Engineer (SRE)

1GLOBAL
Apply
1 month ago
Berlin, GermanySenior

Responsibilities

  • Act as a senior technical contributor within the SRE team, mentor peers, and establish reliability engineering standards.
  • Define, measure, and maintain SLIs and SLOs for infrastructure and customer-facing services.
  • Plan and execute redundancy, resilience, failover, high-availability, disaster recovery, fault-injection, load, and chaos testing.
  • Design automated recovery mechanisms, self-healing workflows, and intelligent alerting systems.
  • Lead incident response, root-cause analysis, blameless post-mortems, and corrective and preventive action tracking.
  • Develop observability using metrics, logs, and traces.
  • Partner with Infrastructure and DevOps teams on deployment safety, rollback policies, and configuration consistency.
  • Reduce operational toil through automation and reliability tooling.
  • Improve on-call practices, alert quality, runbooks, escalation procedures, and incident management processes.
  • Perform capacity planning, performance benchmarking, resilience audits, and cloud cost-optimization initiatives.
  • Maintain internal documentation, playbooks, and operational guidelines.
  • Ensure compliance with security, reliability, and availability standards.

Requirements

  • At least 5 years of experience in Site Reliability, Systems, or Infrastructure Engineering, including 2+ years in a dedicated SRE role.
  • Strong expertise in Linux systems engineering, distributed systems, and networking.
  • Proven experience building and operating highly available, mission-critical production systems.
  • Hands-on experience with redundancy and failover testing, disaster recovery, and high-availability architecture validation.
  • Deep understanding of monitoring, observability, and incident management principles.
  • Experience with Prometheus, Grafana, Loki, Thanos, and OpenTelemetry or similar tools.
  • Proficiency in Python, Go, and Bash for automation and reliability tooling.
  • Strong knowledge of Kubernetes, container orchestration, and service mesh architectures.
  • Experience with AWS, including EKS, EC2, and VPC, plus on-premises infrastructure integration.
  • Proficiency with Infrastructure as Code tools such as Terraform.
  • Understanding of routing, load balancing, BGP, DNS, VXLAN, and other networking fundamentals.
  • Strong analytical, problem-solving, communication, and collaboration skills.
  • Telecom, carrier-grade, or large-scale distributed systems experience is preferred.
  • Experience with chaos engineering, automated failure-scenario validation, capacity planning, traffic engineering, multi-region failover, reliability dashboards, and SRE metrics is preferred.
  • Familiarity with ISO 27001 and NIST SP 800-53 is preferred.

Benefits

  • Hybrid full-time role based in Berlin, Germany.
  • Career growth and professional development alongside industry experts.
  • Opportunities for international experience across 1GLOBAL offices.
  • Collaborative, dynamic, and international work environment with open communication.
  • Exposure to major transactions and the future of the telecommunications industry.
  • Equal opportunity workplace focused on diversity and inclusion.

Tech Stack

Categories

Site Reliability
1GLOBAL

About 1GLOBAL

501-1,000 employees
Contact me