Software Engineer Manager – Platform Reliability Engineering

Posted 2ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Software Engineer Manager leading Reliability Engineering practices for Home Depot's cloud foundation. Ensuring resilience, performance, and security through extensive automation and incident management.

Responsibilities:

  • Ensure the resilience, performance, and security of enterprise cloud foundation
  • Lead a dedicated team of engineers accountable for delivering Reliability Engineering practices
  • Collaborate with product team members (UX, engineering, and product management) to create secure, reliable, scalable software solutions
  • Writes custom code or scripts to automate infrastructure, monitoring services, and test cases
  • Creates meaningful dashboards, logging, alerting, and responses to ensure that issues are captured and addressed proactively
  • Evaluates new technologies for adoption across the enterprise
  • Provides leadership, mentoring, and coaching to Software Engineers
  • Conducts annual and mid-year reviews by reviewing individual development plans and team feedback

Requirements:

  • Must be at least eighteen years of age
  • Must be legally permitted to work in the United States
  • Mastery of an object oriented programming language (preferably Java)
  • 6-10 years of relevant work experience
  • Proven experience managing or leading Site Reliability Engineering (SRE), Platform Engineering, or DevOps teams
  • Demonstrated ability to drive cultural change around process and technology adoption
  • Experience leading incident response, driving mitigation to minimize MTTR, plus blameless post-mortems and RCAs that engineer out recurrence
  • Experience defining and enforcing SLOs/SLIs against Critical User Journeys, using error budgets and Production Readiness Reviews to govern release velocity
  • Strong understanding of Google Cloud or similar public cloud ecosystems
  • Experience operating, troubleshooting, and providing tier-escalation support for complex microservice architectures and high-traffic web applications
  • Strong understanding of modern observability stacks, container orchestration (Kubernetes/GKE), and Infrastructure as Code (Terraform)
  • Strong automation focus: chaos/resiliency testing, toil reduction via custom scripting, and automated recovery/monitoring
  • Solid understanding of software-defined cloud networking, overlay mesh networks, and Zero Trust Network Access (ZTNA) principles
  • Experience leading vulnerability/exposure management with automated remediation SLAs, plus peak-readiness and dynamic-scaling strategies for seasonal surges

Benefits:

  • Health insurance
  • 401(k) matching
  • Professional development opportunities