about 4 hours ago
London, United KingdomStaff+
Responsibilities
- Take ownership of the reliability, capacity, and cost efficiency of DeepL's compute infrastructure.
- Shape the platform's architecture and commit to decisions that serve the best outcomes.
- Design, build, and operate production-grade Kubernetes clusters across cloud and on-prem infrastructure.
- Drive the technical work of deepening the hybrid model for workload unification.
- Define infrastructure-as-code and platform standards, enhancing observability and security.
- Mentor engineers and raise the technical bar through design reviews and architecture decisions.
- Build consensus across engineering teams and implement changes effectively.
- Lead incident response for hybrid infrastructure and sustain an on-call rotation.
Requirements
- Deep, hands-on Kubernetes expertise with experience in designing and operating clusters at scale.
- Experience in either public cloud or on-prem infrastructure, with AWS experience preferred.
- Strong networking and Linux debugging skills.
- Proficiency in infrastructure as code with Terraform or equivalent.
- Software engineering skills in at least one major language, preferably Go or Python.
- A track record of technical ownership across multiple teams.
- Strong incident-management and reliability-engineering discipline.
Benefits
- Diverse and internationally distributed team with over 90 nationalities.
- Open communication and regular feedback culture.
- Hybrid work schedule with flexible hours.
- Virtual Shares for all employees, linking contributions to company growth.
- Regular in-person team events to foster bonding.
- Monthly full-day hacking sessions for personal projects.
- 30 days of annual leave and access to mental health resources.
- Competitive benefits tailored to individual locations.
