
Senior Site Reliability Engineer
PlayStation Global12 hours ago
Adelaide, AustraliaSenior
Responsibilities
- Own production operations, production code quality, deployments, and operational readiness for cloud gaming services.
- Lead improvements in reliability and scalability and define KPIs, processes, and continuous-improvement initiatives with SRE Management.
- Influence architecture and implementation decisions and design platform-wide solutions.
- Mentor junior SRE staff and represent SRE and operational scalability across the wider organization.
- Lead small-scale projects from inception to implementation and provide technical leadership during delivery.
- Participate in an on-call rotation.
Requirements
- At least 7 years of experience in software development and/or Linux systems administration.
- Proficiency as a Linux Production Systems Engineer with experience managing large-scale web services infrastructure.
- Development experience with Python, Bash, Go, Java, C++, or Rust, with Python preferred.
- Experience with at least three relevant areas, including distributed storage, NoSQL, data aggregation, highly available RDBMS, monitoring and alerting, Kubernetes or AWS, software distribution, configuration management, or performance analysis and load testing.
- Strong interpersonal, written, and verbal communication skills.
- Availability to participate in scheduled on-call rotations.
- QA or SDET experience is a plus.
Tech Stack
AnsibleApache CassandraApache HadoopApache KafkaAWSBashC++ChefElasticsearchGoGrafanaJavaKubernetesLinuxMongoDBMySQLPostgreSQLPrometheusPuppetPythonRedisRust
Categories
Site Reliability