Senior Site Reliability Engineer

Runware

Completely RemoteFull TimeInformation Technology
Posted Yesterday

Job description

Responsibilities

  • Own and improve the reliability, availability and performance of critical production services
  • Define and evolve reliability practices including SLIs, SLOs, alerting, and observability
  • Investigate complex production issues across distributed systems, APIs, networking, and GPU-backed workloads
  • Lead incident reviews and RCAs to implement lasting engineering improvements
  • Reduce operational toil through automation and improved deployment safety
  • Partner with Engineering and DevOps on capacity planning and architectural improvements

Requirements

  • Strong experience operating and troubleshooting production systems at scale
  • Deep understanding of distributed systems, databases, queues, and containers
  • Experience designing observability systems using metrics, logs, and distributed tracing
  • Proficiency in SRE principles such as error budgets and capacity planning
  • Experience with Kubernetes, containers, and IaC
  • Ability to write software and automation in Python, Go, or PHP
  • Willingness to participate in an engineering on-call rotation

Preferred Qualifications

  • Experience with high-throughput or low-latency APIs
  • Experience with bare-metal infrastructure, GPU environments, or AI/ML workloads
  • Experience with RabbitMQ or other distributed messaging systems
  • Experience operating MySQL, Redis, or ClickHouse
  • Experience with global traffic management, load balancing, and CDN platforms

Benefits

  • Generous paid time off including vacation, sick days, and public holidays
  • Meaningful stock options
  • Remote-first setup with flexible hours
  • Paid family leave
  • Twice-yearly company retreats in inspiring locations

Skills & tools

KubernetesPythonGo

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 01Experience operating production systems at scale
  2. 02Understanding of distributed systems
  3. 03Experience with Kubernetes and IaC
  4. 04Proficiency in Python, Go, or PHP
  5. 05Experience with observability systems
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches