Vice President Site Reliability Engineering

Galaxy

Completely RemoteFull TimeInformation Technology
Posted Today

Job description

Responsibilities

  • Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets
  • Establish and enforce standards for Infrastructure as Code (IaC) using Terraform
  • Lead strategy for automated configuration and state management using Ansible and Packer
  • Manage monitoring and health of automation platforms using SLIs/SLOs
  • Drive automated lifecycle management for physical and virtual assets
  • Lead development of custom scripts and internal providers using Python, Go, PowerShell, or Bash
  • Collaborate with the Datacenter team to facilitate workflows and system needs
  • Analyze system behavior and resource utilization to optimize automated deployments
  • Provide technical guidance and career mentorship to SREs

Requirements

  • 6-10 years of experience in Infrastructure, SRE, or DevOps focused on automation at scale
  • Deep proficiency with Terraform and Ansible
  • Hands-on experience with image creation using Packer, Ansible, or SCCM
  • Experience managing VMware (vSphere/vCenter) and cloud providers like Azure and AWS
  • High-level scripting skills in Python, Go, PowerShell, and Bash
  • Experience with observability tools such as Splunk, ELK, Prometheus, or Grafana
  • Understanding of network topology and experience with Juniper or Palo Alto
  • Mastery of Git and CI/CD platforms like Jenkins, GitLab CI, or GitHub Actions
  • Proficiency in managing both Windows Server and Linux

Preferred Qualifications

  • Previous experience in team leadership or management
  • Experience with IAM platforms like Entra ID, Active Directory, or Okta
  • Experience with block or object storage (HP Alletra, EMC, DDN, S3, Azure Blob)
  • Experience with storage backup and DR management using Commvault or Veeam

About the Company

Galaxy is a global leader in digital assets and data center infrastructure, delivering solutions that accelerate progress in finance and artificial intelligence.

Skills & tools

TerraformAnsiblePythonAWSSRE

What the team is looking for

Use this list as a quick fit check before you apply.

  1. 016-10 years Infrastructure/SRE/DevOps experience
  2. 02Proficiency in Terraform and Ansible
  3. 03Experience with Packer and image creation
  4. 04Experience with VMware, Azure, and AWS
  5. 05Scripting in Python, Go, PowerShell, or Bash
  6. 06Knowledge of observability tools
  7. 07Understanding of network topology
  8. 08Mastery of Git and CI/CD
  9. 09Windows and Linux administration
NeverApplyAd

Wake up to a shortlist, not a search results page.

NeverApply scores every new listing against your CV, salary floor and visa. A handful of real matches by morning.

Get your daily matches