Vice President, Site Reliability Engineering – Data Centers

Posted 3hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Galaxy VP leading SRE and infrastructure automation across physical and virtual data-center environments. Driving IaC governance, observability, lifecycle management, and custom tooling for digital assets and AI infrastructure.

Responsibilities:

  • Oversee an SRE team focused on designing, deploying, and maintaining automation toolsets and related systems
  • Establish and enforce Infrastructure as Code standards for consistent, repeatable, and secure deployments
  • Lead automated configuration and state management using Ansible playbooks and Packer image pipelines across Windows, Linux, and ESXi platforms
  • Manage monitoring and health of automation platforms and implement SLIs/SLOs
  • Drive automated lifecycle management of physical and virtual assets, including template creation, deployment, patching, scaling, and decommissioning
  • Lead development of custom scripts and internal providers using Python, Go, PowerShell, and Bash
  • Collaborate with the broader Datacenter team and facilitate team-wide workflows
  • Analyze system behavior and resource utilization in virtual environments to optimize automated deployment performance
  • Provide technical guidance and career mentorship to SREs and foster an automate-first culture

Requirements:

  • 6–10 years’ experience in Infrastructure, SRE, or DevOps focused on infrastructure automation at scale
  • Deep proficiency with Terraform, including providers, modules, and state management
  • Deep proficiency with Ansible, including roles, playbooks, and Tower/AWX
  • Hands-on experience creating standardized, hardened Windows and Linux images using Packer, Ansible, or SCCM
  • Strong experience managing and automating VMware vSphere/vCenter, Azure, and AWS
  • High-level scripting skills in Python, Go, PowerShell, and Bash
  • Experience with Splunk, ELK, Prometheus, or Grafana for infrastructure health and automation telemetry
  • Understanding of network topology and design, including Juniper Networks or Palo Alto platforms
  • Strong mastery of Git, including branching strategies and pull-request workflows
  • Experience with CI/CD platforms such as Jenkins, GitLab CI, or GitHub Actions
  • Comfort managing, troubleshooting, and tuning both Windows Server and Linux
  • Previous team leadership or management experience
  • Experience with IAM platforms such as Entra ID, Active Directory, or Okta
  • Experience with block- and object-based storage solutions on-premises or in the cloud
  • Storage backup/disaster recovery administration with Commvault or Veeam

Benefits:

  • Equal employment opportunities
  • Reasonable accommodation for qualified applicants with disabilities