PlayStation Global

Site Reliability Engineer

PlayStation Global
Apply
5 months ago
Berlin, GermanySenior

Responsibilities

  • Lead technical discussions focused on reliability and scalability improvements.
  • Contribute to high-level designs for new products and platforms.
  • Mentor junior SRE staff and support their success.
  • Lead incident response and post-mortem activities for the assigned service team.
  • Collaborate cross-functionally to prioritize reliability improvements, technical-debt reduction, and toil reduction.
  • Contribute code that improves service reliability.
  • Implement automation to reduce ongoing operational toil.
  • Support production ownership, production code quality, deployment, operational readiness, and service stability throughout the software development lifecycle.

Requirements

  • At least 5 years of experience in software development and/or Linux systems administration.
  • Proficiency as a Linux production systems engineer with experience managing large-scale web services infrastructure.
  • Development experience in Python, Bash, Go, Java, C++, or Rust, with Python preferred.
  • Experience with at least three listed areas, including distributed storage, NoSQL, data aggregation, highly available relational databases, monitoring and alerting, Kubernetes or AWS, software distribution, configuration management, or performance analysis and load testing.
  • Strong interpersonal, written, and verbal communication skills.
  • Availability to participate in an on-call rotation.
  • QA or SDET experience is a plus.

Tech Stack

AnsibleApache CassandraApache HadoopApache KafkaAWSBashC++ChefElasticsearchGoGrafanaJavaKubernetesLinuxMongoDBMySQLPostgreSQLPrometheusPuppetPythonRedisRust

Categories

DevOpsSite Reliability
Contact me