Swissquote

Site Reliability Engineer - Cloud Operations

Swissquote
Apply
10 days ago
Gland, SwitzerlandMid Level

Responsibilities

  • Migrate and modernize production applications running on Kubernetes-based platforms.
  • Integrate third-party software into production platforms and align it with operational standards.
  • Design and operate applications on a service mesh platform.
  • Implement canary releases, progressive rollouts, and other safe deployment patterns.
  • Define SLOs, SLIs, and operational KPIs and use them to drive reliability improvements.
  • Improve observability across metrics, logs, and traces.
  • Automate repetitive operational work.
  • Provide Level-3 support and participate in the on-call rotation.
  • Test system behavior under load, during failures, and when dependencies are unavailable.
  • Explore AI tools for troubleshooting, incident analysis, remediation, observability, and operational automation.
  • Operate applications that depend on GPU resources or other AI infrastructure.

Requirements

  • At least 3 years of experience in SRE, DevOps, platform engineering, or a similar production-focused role.
  • Hands-on experience running production workloads on Kubernetes, OpenShift, EKS, or a similar Kubernetes platform.
  • Experience with Helm, service mesh technologies, Kubernetes networking, traffic routing, mTLS, and TLS.
  • Experience with GitOps and deployment strategies such as canary or progressive delivery.
  • Understanding of SRE concepts including SLIs, SLOs, and error budgets.
  • Experience with observability and tracing tools such as Prometheus, Grafana, Elastic Stack, or OpenTelemetry.
  • Strong Linux and networking fundamentals, including TCP/IP, DNS, and load balancing.
  • Ability to troubleshoot JVM-based applications, including heap usage, garbage collection, and JVM configuration issues.
  • Automation experience with Python, Go, Bash, or another programming language.
  • Experience or strong interest in applying AI to observability, incident response, or operational automation.
  • Experience or strong interest in operating applications dependent on GPU resources or AI infrastructure.
  • Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or Puppet.
  • Preferred experience with Argo CD, Argo Rollouts, Argo Workflows, Istio, Linkerd, Envoy, or large-scale Kubernetes platforms.
  • Preferred experience operating Java or Spring Boot applications and tuning JVM applications for performance or low latency.
  • Preferred experience integrating applications with self-hosted AI platforms such as vLLM and troubleshooting GPU availability or NVIDIA MIG configurations.
  • Knowledge of Cilium, eBPF, public cloud, or large private cloud environments is beneficial.
  • CKAD, CKA, CKS, or equivalent hands-on Kubernetes experience is beneficial.
  • Fluent English and good conversational French are required.

Tech Stack

AmbassadorAnsibleArgo CDBashGoGrafanaHelmIstioJavaKubernetesLinuxOpenShiftPrometheusPuppetPythonSpring BootTerraform

Categories

DevOpsSite Reliability
Swissquote

About Swissquote

1,001-5,000 employees
Contact me