Wikimedia Foundation

Senior Site Reliability Engineer, Data Persistence

Wikimedia Foundation
Apply
2 months ago
Remote, WorldwideSenior
H1B Sponsor

Base Salary

$117k - $181k/yr

Responsibilities

  • Perform day-to-day deployment, maintenance, configuration, and troubleshooting of Wikimedia’s public-facing infrastructure.
  • Implement and use configuration-management and deployment tools, including Puppet and Kubernetes.
  • Automate installation, configuration, and maintenance of platform services and lead continuous improvement.
  • Assist product teams with architectural design and operation of scalable services.
  • Participate in a shared 24/7 on-call rotation, including incident response, outage and alert diagnosis, and follow-up.
  • Collaborate asynchronously with a globally distributed, cross-functional team.
  • Mentor peers in areas of technical and operational expertise.
  • Lead and participate in incident reviews, root-cause analysis, and preventive-measure implementation.

Requirements

  • At least 6 years of experience in an SRE, operations, or DevOps role as part of a team.
  • Experience with shell scripting or an SRE scripting language such as Python, Go, Bash, or Ruby, plus configuration-management tools such as Puppet or Ansible.
  • Experience with distributed caching systems, including their algorithms and performance optimization.
  • Experience with package management on Linux systems and strong Linux system-level troubleshooting skills.
  • A history of automating tasks and processes, identifying process gaps, and finding automation opportunities.
  • Strong written and verbal English skills and the ability to work independently across multiple time zones.
  • Experience leading and participating in incident response and post-incident reviews.
  • Preferred experience with Linux kernel tuning; monitoring, metrics, and logging infrastructure; open-source software or communities; LAMP technologies; cross-team SLOs; large-scale filesystems or object stores such as OpenStack Swift or Ceph; distributed storage and database systems such as Cassandra or MariaDB; and SRE backup management with Bacula.
  • MediaWiki experience is a plus.

Benefits

  • Remote-first work arrangement with staff and contractors across more than 40 countries.
  • Travel 1–2 times per year for in-person events and team meetings.
  • U.S. benefits and perks are provided, with hiring available only in the listed U.S. states, territories, and countries.
  • Non-U.S. employees are hired through a local Employer of Record and must have current work authorization in their hiring location.
  • Inclusive equal-opportunity workplace with disability accommodations available during the application process.

Tech Stack

AnsibleApache CassandraBashGoGrafanaKubernetesLinuxMariaDBOpenStackPHPPrometheusPuppetPythonRedisRuby

Categories

DevOpsSite Reliability
Wikimedia Foundation

About Wikimedia Foundation

501-1,000 employees

We are the nonprofit organization that operates Wikipedia and the other Wikimedia free knowledge projects. We host a technology infrastructure that makes possible billions of monthly visits to Wikipedia. Since 2003, we have supported the hundreds of thousands of volunteer editors who edit, expand, and curate the Wikimedia projects. We equip these volunteers with the most up-to-date tools, ensure connections to Wikipedia are fast, safe, and private, and develop new technology and products to meet the demands of our readers and editors. We provide individuals and organizations around the world with funding to increase the knowledge on Wikipedia. We also undertake legal and advocacy efforts to protect people’s right to free knowledge. We are a charitable, not-for-profit organization that relies on donations. We receive donations from millions of individuals around the world, with an average donation of about $11. We also receive donations through institutional grants and gifts. We are a United States 501(c)(3) tax-exempt organization with offices in San Francisco, California.

Contact me