Senior Site Reliability Engineer

Teladoc Health

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

Responsibilities

  • Define, implement, and improve SLIs, SLOs, and error budgets
  • Partner with application teams to improve service reliability and scalability
  • Build and improve observability across applications and infrastructure
  • Analyze system performance, bottlenecks, and capacity risks
  • Support and improve production workloads running on Microsoft Azure
  • Conduct blameless post-incident reviews and document root causes
  • Support operational controls for identity, access, and encryption

Requirements

  • 7+ years in site reliability with ownership of mission-critical services
  • Deep Azure experience including Azure Monitor, Application Insights, and AKS
  • Hands-on experience with enterprise observability platforms like Datadog or Dynatrace

Preferred Qualifications

  • Experience standing up or anchoring an SRE practice
  • Healthcare IT experience with HIPAA or HITRUST familiarity
  • Infrastructure as Code expertise with Terraform or Bicep
  • Scripting proficiency in Python, PowerShell, or Go
  • Multi-cloud experience with AWS

Skills & tools

AzureSREAKSDatadogTerraform

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 017+ years in site reliability
  2. 02Deep Azure experience
  3. 03Azure Monitor
  4. 04Application Insights
  5. 05AKS
  6. 06Cloud-native operations
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches