PlayStation Global

Service Reliability Engineer

PlayStation Global
Apply
6 hours ago
Adelaide, AustraliaSenior

Responsibilities

  • Own production stability, production code quality, deployments, and operational readiness for assigned services.
  • Lead technical discussions focused on reliability, scalability, technical debt, and toil reduction.
  • Contribute to high-level designs for new products and platforms.
  • Mentor junior SRE staff.
  • Lead incident response and post-mortem activities.
  • Contribute code and implement automation to improve reliability and reduce ongoing toil.
  • Participate in an on-call rotation.

Requirements

  • At least 5 years of experience in software development and/or Linux systems administration.
  • Proficiency as a Linux production systems engineer with experience managing large-scale web services infrastructure.
  • Development experience in Python, Bash, Go, Java, C++, or Rust; Python is preferred.
  • Experience with at least three listed areas, including distributed storage, NoSQL, data aggregation, highly available relational databases, monitoring and alerting, Kubernetes or AWS, software distribution, configuration management, or performance analysis and load testing.
  • Strong interpersonal, written, and verbal communication skills.
  • Availability to participate in an on-call rotation.

Tech Stack

AnsibleApache CassandraApache HadoopApache KafkaAWSBashC++ChefElasticsearchGoGrafanaJavaKubernetesLinuxMongoDBMySQLPostgreSQLPrometheusPuppetPythonRedisRust

Categories

Site Reliability
Contact me