2 hours ago
Tokyo, JapanSenior
Responsibilities
- Own the L3 reliability posture for hybrid cloud and on-premises platforms, including SLOs, KPIs, operability gates, production readiness, and runbooks.
- Design and build operational automation, health checks, remediation workflows, and self-service integrations using Terraform, Ansible, and scripting.
- Lead diagnosis, stabilization, recovery, major incident response, root-cause analysis, and preventive actions.
- Define observability standards, actionable signals, dashboards, logging, and alert-quality practices to reduce operational noise.
- Advance predictive and proactive operations through anomaly detection, capacity analysis, and Python-based analytics or ML/DL where applicable.
- Enable and govern outsourced L1/L2 providers through runbooks, training, standard changes, escalation criteria, and ITSM alignment.
- Partner with Platform Engineering on operable-by-design capabilities, operational roadmaps, continuous improvement, and peer mentoring.
Requirements
- 5+ years of experience in platform, SRE, operations, or platform engineering with production ownership in large-scale environments.
- Hands-on hybrid cloud and on-premises operations experience with enterprise fundamentals in compute, networking, storage, and identity.
- Production experience with infrastructure as code and automation using Terraform and Ansible.
- Scripting experience with Python, PowerShell, and/or Bash, with Python strongly preferred.
- Proven L3 incident troubleshooting and major incident leadership experience.
- Strong infrastructure fundamentals, including networking concepts such as DNS and DHCP, virtualization, storage, Windows Server, and/or Linux.
- Experience with ITSM processes including incident, problem, and change management and ticket-based operations.
- Knowledge of the Azure platform.
- Preferred qualifications include SRE practices, virtualization, backup and disaster recovery, containers, configuration drift control, operational analytics, ML/DL for anomaly detection or forecasting, large multinational 24/7 operations, and AI or agentic automation approaches.
- Strong ownership, communication, documentation, incident leadership, cross-team collaboration, and vendor-partnering skills.
Benefits
- Elective benefits tailored to the employee’s country and lifestyle.
- Formal leadership and professional development programs, on-demand courses, and career-growth resources.
- Financial, physical, and mental well-being support through seminars, events, and a global Life Empowerment Assistance Program.
- Inclusive education, peer-to-peer conversations, diversity and inclusion initiatives, and equitable growth opportunities.
- Global onboarding and networking with new coworkers within the first 30 days.
- Internal peer-led communities, business resource groups, volunteering, and environmental and social initiatives.
