Lunit, Inc.

(Seoul) Senior Site Reliability Engineer· Cancer Screening

Lunit, Inc.
Apply
20 hours ago
Seoul, Korea, SouthSenior

Responsibilities

  • Improve the stability, availability, performance, and operational quality of the Lunit INSIGHT product and related services.
  • Design and enhance cloud infrastructure, deployment, monitoring, observability, and operational automation systems.
  • Analyze cloud resource usage and cost and drive cost optimization while maintaining stability and performance.
  • Diagnose, recover from, and perform root-cause analysis on production incidents, then prevent recurrence through monitoring, alerting, runbooks, and automation.
  • Manage and improve operational configurations and security elements including environment and customer settings, certificates, secrets, and access permissions.
  • Collaborate with product and development teams to build reliability into product design and development.
  • Coordinate in English with Lunit International's SRE team and align operational environments and processes.
  • Build and participate in a 24/7 on-call and incident-response system and evolve a sustainable on-call operating model.

Requirements

  • 5+ years of experience in SRE, DevOps, Platform Engineering, or production infrastructure operations.
  • Direct experience designing and operating Azure-based production environments.
  • Hands-on experience with Linux, networking, and containers.
  • Experience designing or improving CI/CD, Infrastructure as Code, or operational automation systems.
  • Experience analyzing production incidents using monitoring and performing recovery and recurrence-prevention activities.
  • Experience collaborating with product and development teams to balance stability and development velocity.
  • Fluent English communication for technical decisions and coordination with overseas engineers and colleagues from diverse roles.
  • Preferred experience establishing technical direction or operating systems for SRE or Platform domains.
  • Preferred experience leading structural improvements for recurring incidents or operational inefficiencies, including recurrence prevention and automation.
  • Preferred experience with on-call, incident management, postmortems, and SRE practices.
  • Preferred experience with Terraform, Bicep, ARM templates, Python, and Bash for Infrastructure as Code and operational automation.
  • Preferred experience operating Kubernetes, GitHub Actions, and Azure DevOps deployment environments.
  • Preferred experience with Azure Monitor, Application Insights, Log Analytics, PagerDuty, and ServiceNow.

Benefits

  • Full-time position with a 3-month probation period.
  • On-site work at Lunit headquarters near Gangnam Station in Seoul.
  • Meal allowance of up to 12,000 KRW per meal when working at the office.
  • Latest computers such as Macs and 4K monitors, renewable every three years.
  • Seminar registration fees and book purchases covered.
  • Regular in-house AI and medical seminars and access to AI learning resources and a deep learning DevOps system.
  • In-house English lessons provided.
  • Up to 1.2 million KRW in annual benefits points.
  • Holiday gifts or vouchers for Korean National holidays, Seollal, and Chuseok.
  • Congratulatory and condolence allowances with paid time off.
  • Annual medical checkups and employee accident insurance.

Tech Stack

AzureBashDockerFastAPIGitHub ActionsGoKubernetesLinuxPostgreSQLPythonRabbitMQRedisTerraform

Categories

DevOpsSite Reliability
Lunit, Inc.

About Lunit, Inc.

201-500 employees
Contact me