Site Reliability Engineer

Outpost

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

About the Company

Outpost is building the backbone of freight, reinventing supply chain infrastructure with carrier-agnostic truck terminals. As a vertically integrated real estate, operations, and technology company, Outpost acquires and operates mission-critical real estate to serve the world's largest logistics providers.

Responsibilities

  • Own reliability targets across backend/API, worker services, applications, and CV pipelines
  • Level up monitoring and alerting, and build out auto-remediation systems
  • Partner with engineering to build agents that triage alerts and handle routine remediation
  • Harden and optimize GCP infrastructure (Cloud Run, Cloud SQL, GCS) for cost and performance
  • Own database scale and performance, including connection pooling, query optimization, and capacity planning
  • Improve the reliability of ML training and monitoring infrastructure
  • Run blameless postmortems and drive fixes for root causes
  • Participate in on-call rotation

Requirements

  • 4+ years in an SRE, infrastructure, or backend engineering role with production on-call ownership
  • Deep experience with a major cloud provider, preferably GCP
  • Experience building monitoring, alerting, and observability stacks (Grafana, Prometheus, Datadog, etc.)
  • Strong scripting and automation skills in Python or Bash
  • Proficiency with containerized workloads (Docker) and CI/CD pipelines
  • Proven track record of reducing incident volume or improving reliability metrics
  • Strong English communication skills for clear writing and asynchronous engagement

Preferred Qualifications

  • Experience with ML/data infrastructure, training pipelines, or model monitoring
  • Experience building or integrating AI agents for operational automation
  • Infrastructure-as-code experience with Terraform
  • PostgreSQL performance tuning at scale
  • Background supporting physical/IoT systems or edge devices
  • Experience with bare-metal infrastructure and colocation environments

Skills & tools

GCPDockerPythonPostgreSQLTerraform

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 014+ years SRE or infrastructure experience
  2. 02Deep GCP experience
  3. 03Monitoring/observability stack experience
  4. 04Python or Bash scripting
  5. 05Docker and CI/CD experience
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches