
Staff SRE Engineer
Stellar Cyber4 months ago
Remote, SpainStaff+
Responsibilities
- Administer container orchestration platforms and containerized workloads.
- Monitor and troubleshoot production systems and participate in on-call rotations.
- Improve monitoring, logging, alerting, and observability across systems and data platforms.
- Administer and optimize cloud environments across multiple providers.
- Manage distributed data platforms and real-time processing systems.
- Develop and maintain CI/CD pipelines and Infrastructure as Code practices.
- Automate and orchestrate infrastructure using programming and scripting languages.
- Perform system administration and networking tasks.
- Lead or influence incident response, on-call practices, reliability culture, architecture, tooling, and operational best practices.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering roles.
- Proven success operating large-scale production systems in AWS, GCP, Azure, or OCI cloud environments.
- Demonstrated leadership in incident response, on-call best practices, and reliability-focused practices.
- Advanced Kubernetes administration and troubleshooting experience.
- Hands-on experience with Prometheus, Grafana, Loki, and Alertmanager.
- Knowledge of chat-based operations interfaces or AI-agent-based auto-remediation controllers, including alert triage and root-cause hypothesis generation.
- Experience operating Elasticsearch, MongoDB, Spark, Kafka, and Redis.
- Strong Python and Bash programming and automation skills.
- Deep understanding of Terraform, Helm, CI/CD pipelines, distributed systems, databases, networking, and Linux administration.
- Bachelor's degree in Computer Science, Engineering, or a related technical field.
- AWS, GCP, observability, Linux, or Kubernetes certifications are preferred.
Benefits
- Remote work arrangement in Spain.
- Full-time position.
- The interview process may include an in-person interview.
Tech Stack
Apache KafkaApache SparkAWSAzureBashElasticsearchGitHub ActionsGoogle Cloud PlatformGrafanaHelmKubernetesLinuxMongoDBPrometheusPythonRedisTerraform
Categories
Site Reliability