Site Reliability Engineer - APAC
Tyk Technologies6 months ago
Remote, Worldwide or Hong Kong, Hong KongSenior
Responsibilities
- Maintain the global Tyk Cloud platform within defined service objectives.
- Identify and resolve reliability issues with the squad.
- Define metrics and build operational dashboards.
- Participate in the on-call rotation and manage incidents.
- Expand the platform’s multi-region and multi-cloud capabilities.
- Document operational knowledge, SRE processes, and policies.
- Conduct post-incident analysis and automate common operational tasks.
- Recommend and implement ways to improve operational efficiency and reduce operating costs without affecting service.
- Assist with cloud penetration testing through provider coordination, technical details, and environment setup.
Requirements
- Advanced Kubernetes, containers, AWS/EKS, and Linux skills.
- Experience launching and operating production-scale Kubernetes clusters.
- Experience designing and operating infrastructure on AWS and other providers.
- Experience operating MongoDB or other document database clusters and Redis or other key-value storage clusters.
- Experience maintaining distributed software and operating Prometheus and Grafana.
- Proficiency with Terraform and infrastructure as code and with Helm.
- Familiarity with Thanos, networking concepts, and DNS, TCP/IP, HTTP, TLS, and UDP protocols.
- Strong collaboration skills and a proactive, energetic, innovative, and change-oriented approach.
- Nice-to-have experience with GCP or Azure, bare-metal infrastructure, API management, large-scale distributed storage, Rancher, or Go.
- CKA, CKAD, or CKS certification is a nice-to-have.
Benefits
- Unlimited paid holidays.
- Remote working from anywhere in the world with flexible working hours.
- Employee share scheme.
- Generous maternity and paternity leave.
- Volunteering days.
- Employee wellbeing platform.
- Full-time role with an on-call rotation from 16:00 to 04:00 UTC.
Tech Stack
Categories
Site Reliability