
Site Reliability Engineer - Cloud Operations
Swissquote10 days ago
Gland, SwitzerlandMid Level
Responsibilities
- Migrate and modernize production applications running on Kubernetes-based platforms.
- Integrate third-party software into production platforms and align it with operational standards.
- Design and operate applications on a service mesh platform.
- Implement canary releases, progressive rollouts, and other safe deployment patterns.
- Define SLOs, SLIs, and operational KPIs and use them to drive reliability improvements.
- Improve observability across metrics, logs, and traces.
- Automate repetitive operational work.
- Provide Level-3 support and participate in the on-call rotation.
- Test system behavior under load, during failures, and when dependencies are unavailable.
- Explore AI tools for troubleshooting, incident analysis, remediation, observability, and operational automation.
- Operate applications that depend on GPU resources or other AI infrastructure.
Requirements
- At least 3 years of experience in SRE, DevOps, platform engineering, or a similar production-focused role.
- Hands-on experience running production workloads on Kubernetes, OpenShift, EKS, or a similar Kubernetes platform.
- Experience with Helm, service mesh technologies, Kubernetes networking, traffic routing, mTLS, and TLS.
- Experience with GitOps and deployment strategies such as canary or progressive delivery.
- Understanding of SRE concepts including SLIs, SLOs, and error budgets.
- Experience with observability and tracing tools such as Prometheus, Grafana, Elastic Stack, or OpenTelemetry.
- Strong Linux and networking fundamentals, including TCP/IP, DNS, and load balancing.
- Ability to troubleshoot JVM-based applications, including heap usage, garbage collection, and JVM configuration issues.
- Automation experience with Python, Go, Bash, or another programming language.
- Experience or strong interest in applying AI to observability, incident response, or operational automation.
- Experience or strong interest in operating applications dependent on GPU resources or AI infrastructure.
- Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or Puppet.
- Preferred experience with Argo CD, Argo Rollouts, Argo Workflows, Istio, Linkerd, Envoy, or large-scale Kubernetes platforms.
- Preferred experience operating Java or Spring Boot applications and tuning JVM applications for performance or low latency.
- Preferred experience integrating applications with self-hosted AI platforms such as vLLM and troubleshooting GPU availability or NVIDIA MIG configurations.
- Knowledge of Cilium, eBPF, public cloud, or large private cloud environments is beneficial.
- CKAD, CKA, CKS, or equivalent hands-on Kubernetes experience is beneficial.
- Fluent English and good conversational French are required.
Tech Stack
Categories
DevOpsSite Reliability