Senior Site Reliability Engineer

Wikimedia Foundation

Completely RemoteFull TimeInformation Technology
Posted 3 days ago

Job description

Responsibilities

  • Perform day-to-day operational and DevOps tasks on public-facing infrastructure including deployment, maintenance, and troubleshooting
  • Implement and utilize configuration management and deployment tools like Puppet and Kubernetes
  • Automate the installation, configuration, and maintenance of services to drive continuous improvement
  • Assist product teams with architectural design to ensure services are scalable
  • Participate in a 24/7 on-call rotation for incident response and diagnosis
  • Collaborate with a global, cross-functional team in an asynchronous environment
  • Mentor peers in technical and operational areas

Requirements

  • 6+ years of experience in an SRE, Operations, or DevOps role
  • Proficiency in shell scripting and languages such as Python, Go, Bash, or Ruby
  • Experience with configuration management tools like Puppet or Ansible
  • Experience with distributed caching systems and performance optimization
  • Experience with Linux package management (Debian preferred)
  • Strong Linux system-level troubleshooting skills
  • Proven history of automating tasks and identifying process gaps
  • Strong English verbal and written communication skills
  • Experience leading incident response and conducting root cause analysis

Preferred Qualifications

  • Experience with Linux kernel tuning
  • Experience with monitoring and logging infrastructure such as Prometheus and Grafana
  • Contributions to Free and Open Source software
  • Experience with LAMP stack technologies (PHP/HHVM, memcached/Redis)
  • Experience operating on-premise filesystems or object stores like OpenStack Swift or Ceph
  • Experience with distributed storage and database systems like Cassandra or MariaDB
  • Experience managing backups with tools like Bacula

About the Company

The Wikimedia Foundation is the nonprofit organization that operates Wikipedia and other free knowledge projects. Our vision is a world in which every single human can freely share in the sum of all knowledge.

Skills & tools

PythonKubernetesPuppet

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 016+ years SRE/DevOps experience
  2. 02Python, Go, Bash, or Ruby
  3. 03Puppet or Ansible
  4. 04Distributed caching systems
  5. 05Linux system troubleshooting
  6. 06Incident response experience
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches