8 days ago
Seoul, Korea, SouthSenior
Responsibilities
- Develop AIOps platform transformation strategies and provide technical advisory support.
- Design AIOps platform and observability pipeline architectures.
- Develop scale-up strategies for auto-healing and auto-scaling operations.
- Research automation feasibility and analyze return on investment.
- Define SLI, SLO, and SLA practices and lead organization-wide SRE culture adoption.
- Build and operate automated runbook execution systems and platform tools.
Requirements
- At least 7 years of practical experience in SRE, DevOps, or Platform Engineering.
- Experience independently defining SRE culture and methodologies and introducing them across an organization.
- Experience building and operating systems in cloud Kubernetes environments and building GitOps and Argo CD-based delivery pipelines.
- Experience building and operating observability stacks using multiple technologies such as Prometheus, Grafana, ELK, and OpenTelemetry.
- Experience responding to infrastructure incidents through on-call operations and improving systems through post-mortems.
- Experience with infrastructure automation tools including Terraform, Karpenter, and Ansible.
- No restrictions preventing overseas business travel.
- Experience introducing AIOps or deployment automation into production services is preferred.
- Experience using AI coding agents such as Claude Code CLI or Codex CLI for automation is preferred.
- Experience developing platform tools such as MCP, Skills, or Plugins to improve developer productivity is preferred.
- Experience applying SLI, SLO, and Error Budget concepts from SRE practices such as the Google SRE Book is preferred.
Benefits
- Professional contract employment with a 5-month probationary period.
- Work location: Yeoksam Centerfield East Tower.
- The position is open until filled and may close early; interviews and job tests may be added as needed.
- Hiring process includes resume screening, phone interview, personality assessment, technical fit interview, and two to three culture fit interviews.
Tech Stack
Categories
DevOpsSite Reliability
