
Site Reliability Engineer - eFX/ Crypto
Swissquote12 days ago
Gland, SwitzerlandMid Level
Responsibilities
- Monitor production systems and respond to incidents to maintain uptime, performance, and availability.
- Operate, maintain, and troubleshoot Kubernetes clusters, containerized applications, and on-premise infrastructure.
- Design and implement scalable, reliable infrastructure solutions across on-premise and private cloud environments.
- Automate IT and operational processes and maintain CI/CD pipelines and deployment automation.
- Implement and improve observability and monitoring using Prometheus, Grafana, and Elastic/Kibana.
- Perform fault analysis, root-cause analysis, corrective actions, preventive actions, stress testing, resilience testing, disaster-recovery testing, and BCP testing.
- Collaborate with Infrastructure and Development teams to improve production readiness and reliability.
- Participate in the production on-call rotation and continuously improve operational processes, monitoring, automation, and reliability practices.
Requirements
- At least 3 years of experience in SRE, DevOps, Systems Engineering, Platform Engineering, or a similar production engineering role.
- Mandatory hands-on Kubernetes experience in production, including Kubernetes architecture, service mesh and Istio, user management, health checks and probes, networking and DNS, troubleshooting, and performance analysis.
- Linux administration and troubleshooting skills, plus experience with Docker and container technologies.
- Experience with on-premise infrastructure and/or private cloud environments.
- Experience with Infrastructure as Code and automation tools such as Terraform and Ansible, and with CI/CD pipelines.
- Experience with observability and monitoring tools such as Prometheus, Grafana, and Elastic/Kibana.
- Good programming and scripting skills in Python and Shell.
- Strong networking fundamentals, including TCP/IP, DNS, HTTPS, HTTP, and load balancing.
- Understanding of SRE principles including SLIs, SLOs, SLAs, availability, reliability, and service performance.
- Understanding of Java web applications and web servers such as Spring Boot and Apache Tomcat.
- Fluent written and spoken English and willingness to learn French.
- Nice-to-have qualifications include CKA or CKAD certification; Helm, ArgoCD, GitHub Actions, or Kubernetes operator experience; Apache Kafka or RabbitMQ knowledge; MT4/MT5 knowledge; Windows Server 2016/2022 experience; financial-services experience; and experience with highly available or low-latency production systems.
Benefits
- The employer offers an equal-opportunity workplace welcoming candidates from all backgrounds, experiences, and perspectives.
- The company describes opportunities to try new things, own work, and grow in a fast-growing digital banking environment.
- The workplace has no dress code and promotes team celebrations and an informal atmosphere.
Tech Stack
AnsibleApache KafkaDockerGitHub ActionsGrafanaHelmIstioKibanaKubernetesLinuxPrometheusPythonRabbitMQSpring BootTerraform
Categories
Site Reliability