LivePerson

Senior Site Reliability Engineer

LivePerson
Apply
18 hours ago
Sofia, BulgariaSenior

Responsibilities

  • Design, build, and maintain highly available, scalable, secure, and resilient infrastructure and services across cloud and hybrid environments, primarily on Google Cloud Platform.
  • Develop automation and infrastructure-as-code solutions using Python, Terraform, Ansible, Bash, and related SRE tools.
  • Deploy, operate, and troubleshoot Kubernetes-based platforms and workloads across production environments.
  • Implement GitOps deployment workflows using Kubernetes, Helm, and FluxCD.
  • Build and maintain GitLab CI/CD pipelines for application and infrastructure delivery.
  • Establish observability practices using metrics, logs, traces, dashboards, and alerting.
  • Define and improve SLOs, SLIs, and reliability metrics for critical services.
  • Lead or participate in incident response, root cause analysis, corrective actions, and post-incident reviews.
  • Drive automation, capacity planning, performance analysis, reliability assessments, architecture improvements, and operational process improvements.
  • Collaborate with engineering, security, networking, and infrastructure teams and mentor other engineers.
  • Participate in an on-call rotation supporting critical production services.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • At least 5 years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, systems engineering, or a related field.
  • Strong programming and scripting experience with Python, Bash, or similar languages focused on automation and operational tooling.
  • Hands-on experience with Google Cloud Platform, cloud infrastructure, networking, IAM, compute, storage, and managed services.
  • Extensive experience with Kubernetes, Docker, Terraform, Ansible, Helm, FluxCD, and GitLab CI/CD.
  • Understanding of Linux systems administration, networking, DNS, TLS/SSL, authentication, and security fundamentals.
  • Experience with monitoring and observability platforms such as Prometheus, Grafana, and Alertmanager.
  • Experience designing and operating highly available distributed systems and managing production incidents.
  • Strong troubleshooting, problem-solving, communication, collaboration, ownership, and cross-team influence skills.
  • Experience mentoring engineers and influencing technical decisions across teams.
  • Preferred experience with large-scale B2B SaaS or distributed production environments, Istio, HashiCorp Vault, hybrid or on-premises infrastructure, enterprise networking, and internal engineering platforms or self-service tooling.

Benefits

  • Medical, dental, and vision benefits.
  • Native AI learning and professional development opportunities.
  • Employee resource groups and an inclusive, equal-opportunity workplace.
  • The role includes participation in an on-call rotation; work arrangement and contract duration are not stated.

Tech Stack

AnsibleBashDockerGitLab CI/CDGoogle Cloud PlatformGrafanaHelmIstioKubernetesLinuxPrometheusPythonTerraform

Categories

Site Reliability
LivePerson

About LivePerson

1,001-5,000 employees

LivePerson builds conversational AI and messaging software for enterprises, delivered via its Conversational Cloud, to power customer service, commerce, and engagement across web, mobile, and voice channels. The company monetizes through SaaS subscriptions and services, and lists brands such as HSBC, Chipotle, and Virgin Media as customers. Founded in 1995 and headquartered in New York City, LivePerson is a public company traded on NASDAQ under the ticker LPSN.

Contact me