Senior Site Reliability Engineer, Data Persistence

Wikimedia Foundation

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

Responsibilities

  • Perform day-to-day operational/DevOps tasks on public facing infrastructure
  • Implement and utilize configuration management and deployment tools like Puppet and Kubernetes
  • Automate the installation, configuration, and maintenance of services
  • Assist product teams in architectural design for scalable functionality
  • Participate in a 24/7 on-call rotation and incident response
  • Collaborate with a global, cross-functional team asynchronously
  • Mentor peers in technical and operational areas

Requirements

  • 6+ years of experience in an SRE, Operations, or DevOps role
  • Proficiency in shell scripting and languages like Python, Go, Bash, or Ruby
  • Experience with configuration management tools such as Puppet or Ansible
  • Experience with distributed caching systems and performance optimization
  • Experience with Linux package management (Debian preferred)
  • Strong Linux system-level troubleshooting skills
  • Proven history of automating tasks and identifying process gaps
  • Strong English verbal and written communication skills
  • Experience leading incident response and conducting root cause analysis

Preferred Qualifications

  • Experience with Linux kernel tuning
  • Experience with monitoring, metrics, and logging infrastructure (Prometheus, Grafana)
  • Contributions to Free and Open Source software
  • Experience with LAMP stack technologies (PHP/HHVM, memcached/Redis)
  • Experience operating on-premise filesystems or object stores (OpenStack Swift, Ceph)
  • Experience with distributed storage and database systems (Cassandra, MariaDB)
  • Experience managing backups with Bacula

About the Company

The Wikimedia Foundation is the nonprofit organization that operates Wikipedia and other free knowledge projects. Our vision is a world in which every single human can freely share in the sum of all knowledge.

Skills & tools

PythonKubernetesPuppetLinuxDevOps

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 016+ years SRE/DevOps experience
  2. 02Python, Go, Bash, or Ruby proficiency
  3. 03Puppet or Ansible experience
  4. 04Distributed caching systems knowledge
  5. 05Linux system troubleshooting
  6. 06Incident response experience
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches