
Senior SRE Engineer
Stellar Cyber4 months ago
Remote, SpainSenior
Responsibilities
- Administer container orchestration platforms and containerized workloads.
- Monitor and troubleshoot production systems and participate in on-call rotations.
- Improve monitoring, logging, alerting, and observability across systems and data platforms.
- Administer and optimize cloud environments across multiple providers.
- Manage distributed data platforms and real-time processing systems.
- Develop and maintain continuous integration and delivery pipelines.
- Own and implement Infrastructure as Code practices.
- Automate and orchestrate infrastructure using programming and scripting languages.
- Perform systems administration and networking tasks.
- Drive incident response, on-call best practices, and reliability-focused operational improvements.
Requirements
- At least 5 years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
- Proven success operating large-scale production systems in cloud environments including AWS, GCP, Azure, or OCI.
- Demonstrated leadership in incident response, on-call practices, and reliability culture.
- Advanced Kubernetes administration and troubleshooting experience.
- Hands-on experience with Prometheus, Grafana, Loki, and Alertmanager.
- Knowledge of chat-based operations interfaces or AI-agent-based auto-remediation controllers.
- Experience operating Elasticsearch, MongoDB, Spark, Kafka, and Redis.
- Strong Python and Bash programming and automation skills.
- Deep understanding of Terraform and Helm for Infrastructure as Code.
- Experience with GitHub Actions, Bitbucket, and ArgoCD pipelines.
- Strong background in distributed systems, databases, networking, and Linux administration.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- AWS, GCP, observability, Linux, or Kubernetes certifications are preferred.
Benefits
- Full-time remote position based in Spain.
- The interview process may include an in-person interview with the team.
Tech Stack
Apache KafkaApache SparkAWSAzureBashElasticsearchGitHub ActionsGoogle Cloud PlatformGrafanaHelmKubernetesLinuxMongoDBPrometheusPythonRedisTerraform
Categories
Site Reliability