3 hours ago
Berlin, Germany or Munich, GermanyStaff+
Responsibilities
- Own the reliability, capacity, cost efficiency, and hardware lifecycle of CPU and GPU compute infrastructure across AWS and on-premises environments.
- Shape platform architecture and make technical decisions across engineering teams.
- Design, build, and operate Kubernetes clusters through their full lifecycle in cloud and on-premises environments.
- Drive unification of workloads across on-premises hardware and AWS and support cloud-native adoption by research teams.
- Define infrastructure-as-code and platform standards while improving observability and security.
- Mentor infrastructure engineers and raise the technical quality of the group through design reviews and architecture decisions.
- Build cross-team consensus and turn platform direction into shipped changes.
- Lead hybrid-infrastructure incident response, drive root-cause remediation, and help maintain a sustainable on-call rotation.
Requirements
- Deep hands-on experience designing, building, and operating Kubernetes clusters at scale.
- Extensive experience with public cloud or on-premises infrastructure, with AWS preferred and bare-metal or data-center experience valued.
- Strong networking and Linux debugging skills across containers, hosts, and network edges.
- Production experience with Terraform or equivalent infrastructure-as-code tools and GitOps delivery such as ArgoCD.
- Software engineering experience in at least one major language, with Go or Python preferred, including building tools used by other engineers.
- Demonstrated technical ownership across multiple teams and the ability to drive decisions across organizational boundaries.
- Strong incident-management and reliability-engineering discipline.
- Bonus: experience with GPU infrastructure for training or inference, production distributed storage such as Ceph, platforms serving researchers or data scientists, or BGP networking in hybrid or on-premises environments.
Benefits
- Hybrid work with office attendance twice per week and flexible working hours.
- Virtual Shares for every employee.
- Regular in-person team and company events.
- Monthly full-day Hack Friday sessions.
- 30 days of annual leave excluding public holidays, plus mental health resources.
- Competitive location-tailored benefits.
About DeepL
DeepL is a global communications platform powered by Language AI. Since 2017, we’ve been on a mission to break down language barriers. Our human-sounding translations and intelligent writing suggestions are designed with enterprise security in mind. Today, they enable over 200,000 businesses to transform communications, reach new markets, and improve productivity. And, empower millions of individuals around the world to make sense of the world and express their ideas. Join us in exploring the possibilities of Language AI!
