
(Seoul) Senior Site Reliability Engineer· Cancer Screening
Lunit, Inc.20 hours ago
Seoul, Korea, SouthSenior
Responsibilities
- Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related services.
- Design and enhance cloud infrastructure, deployment, monitoring, observability, and operational automation systems.
- Analyze cloud resource usage and cost and drive cost optimization while maintaining stability and performance.
- Diagnose, recover from, and perform root-cause analysis on production incidents, then prevent recurrence through monitoring, alerting, runbooks, and automation.
- Manage and improve operational configurations and security elements including environment and customer settings, certificates, secrets, and access permissions.
- Collaborate with product and development teams to build reliability into product design and development.
- Coordinate in English with Lunit International's SRE team and align operational environments and processes.
- Build and participate in a 24/7 on-call and incident-response system and evolve a sustainable on-call operating model.
Requirements
- 5+ years of experience in SRE, DevOps, Platform Engineering, or production infrastructure operations.
- Direct experience designing and operating Azure-based production environments.
- Hands-on experience with Linux, networking, and containers.
- Experience designing or improving CI/CD, Infrastructure as Code, or operational automation systems.
- Experience analyzing production incidents using monitoring and performing recovery and recurrence-prevention activities.
- Experience collaborating with product and development teams to balance stability and development velocity.
- Fluent English communication for technical decisions and coordination with overseas engineers and colleagues from diverse roles.
- Preferred experience establishing technical direction or operating systems for SRE or Platform domains.
- Preferred experience leading structural improvements for recurring incidents or operational inefficiencies, including recurrence prevention and automation.
- Preferred experience with on-call, incident management, postmortems, and SRE practices.
- Preferred experience with Terraform, Bicep, ARM templates, Python, and Bash for Infrastructure as Code and operational automation.
- Preferred experience operating Kubernetes, GitHub Actions, and Azure DevOps deployment environments.
- Preferred experience with Azure Monitor, Application Insights, Log Analytics, PagerDuty, and ServiceNow.
Benefits
- Full-time position with a 3-month probation period.
- On-site work at Lunit headquarters near Gangnam Station in Seoul.
- Meal allowance of up to 12,000 KRW per meal when working at the office.
- Latest computers such as Macs and 4K monitors, renewable every three years.
- Seminar registration fees and book purchases covered.
- Regular in-house AI and medical seminars and access to AI learning resources and a deep learning DevOps system.
- In-house English lessons provided.
- Up to 1.2 million KRW in annual benefits points.
- Holiday gifts or vouchers for Korean National holidays, Seollal, and Chuseok.
- Congratulatory and condolence allowances with paid time off.
- Annual medical checkups and employee accident insurance.
Tech Stack
Categories
DevOpsSite Reliability