1 month ago
Berlin, GermanySenior
Responsibilities
- Act as a senior technical contributor within the SRE team, mentor peers, and establish reliability engineering standards.
- Define, measure, and maintain SLIs and SLOs for infrastructure and customer-facing services.
- Plan and execute redundancy, resilience, failover, high-availability, disaster recovery, fault-injection, load, and chaos testing.
- Design automated recovery mechanisms, self-healing workflows, and intelligent alerting systems.
- Lead incident response, root-cause analysis, blameless post-mortems, and corrective and preventive action tracking.
- Develop observability using metrics, logs, and traces.
- Partner with Infrastructure and DevOps teams on deployment safety, rollback policies, and configuration consistency.
- Reduce operational toil through automation and reliability tooling.
- Improve on-call practices, alert quality, runbooks, escalation procedures, and incident management processes.
- Perform capacity planning, performance benchmarking, resilience audits, and cloud cost-optimization initiatives.
- Maintain internal documentation, playbooks, and operational guidelines.
- Ensure compliance with security, reliability, and availability standards.
Requirements
- At least 5 years of experience in Site Reliability, Systems, or Infrastructure Engineering, including 2+ years in a dedicated SRE role.
- Strong expertise in Linux systems engineering, distributed systems, and networking.
- Proven experience building and operating highly available, mission-critical production systems.
- Hands-on experience with redundancy and failover testing, disaster recovery, and high-availability architecture validation.
- Deep understanding of monitoring, observability, and incident management principles.
- Experience with Prometheus, Grafana, Loki, Thanos, and OpenTelemetry or similar tools.
- Proficiency in Python, Go, and Bash for automation and reliability tooling.
- Strong knowledge of Kubernetes, container orchestration, and service mesh architectures.
- Experience with AWS, including EKS, EC2, and VPC, plus on-premises infrastructure integration.
- Proficiency with Infrastructure as Code tools such as Terraform.
- Understanding of routing, load balancing, BGP, DNS, VXLAN, and other networking fundamentals.
- Strong analytical, problem-solving, communication, and collaboration skills.
- Telecom, carrier-grade, or large-scale distributed systems experience is preferred.
- Experience with chaos engineering, automated failure-scenario validation, capacity planning, traffic engineering, multi-region failover, reliability dashboards, and SRE metrics is preferred.
- Familiarity with ISO 27001 and NIST SP 800-53 is preferred.
Benefits
- Hybrid full-time role based in Berlin, Germany.
- Career growth and professional development alongside industry experts.
- Opportunities for international experience across 1GLOBAL offices.
- Collaborative, dynamic, and international work environment with open communication.
- Exposure to major transactions and the future of the telecommunications industry.
- Equal opportunity workplace focused on diversity and inclusion.
