28 days ago
Remote, EMEA or London, United KingdomSenior
Responsibilities
- Co-own production services and ensure they are built, maintained, and operated reliably and scalably.
- Support delivery of new features and services as well as day-to-day operations of existing services.
- Develop performance benchmarks and analyze metrics to identify operational improvements.
- Build tooling and automation, maintain monitoring and CI/CD systems, and improve deployment workflows.
- Manage and scale AWS and application infrastructure, including containerized production environments.
- Implement and improve blue/green and canary deployment solutions.
- Participate in a weekly on-call rotation to investigate and resolve system issues.
- Collaborate with Software Engineering to optimize local development workflows and reduce recurring operational issues.
Requirements
- Extensive experience deploying, managing, and troubleshooting infrastructure in AWS.
- Experience managing the full lifecycle of production containers using self-managed Kubernetes, ECS, or EKS.
- Ability to build custom tools when suitable tools are unavailable and teach others to use them.
- Experience building custom tooling to deploy code to production environments.
- Ability to troubleshoot distributed Linux systems and trace requests across applications, systems, and networks.
- Ability to automate routine tasks and proficiency in at least two programming languages.
- Strong written and spoken communication skills for explaining complex ideas to varied audiences.
- Ability to work collaboratively and independently.
Benefits
- Opportunity to earn equity.
- Maternity and paternity leave.
- WeWork membership.
- Work-from-home yearly stipend.
- Learning and development stipend after six months.
- Candidates must be based in the EMEA region and available during standard EMEA working hours.
Tech Stack
Categories
Site Reliability
