Head of Cloud Operations

Posted 3hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

HavocAI head leading cloud operations for defense and commercial autonomous systems. Directing SRE, DevOps, reliability, releases, incidents, and compliance.

Responsibilities:

  • Lead and manage the SRE and DevOps teams, setting technical and operational direction across reliability, automation, and safe delivery.
  • Hire, coach, develop, and manage performance for engineers across both functions.
  • Own reliability and delivery practices including SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.
  • Establish clear ownership and operating expectations across cloud reliability and delivery.
  • Partner with engineering leaders to identify systemic reliability risks and prioritize improvements.
  • Define and own change classes, approval paths, impact assessments, rollback plans, and approval records.
  • Integrate change management with GitOps workflows using merged, signed, peer-reviewed pull requests.
  • Own maintenance windows, freeze periods, emergency-change processes, release calendar, versioning strategy, and promotion across environments and tenants.
  • Establish pre-deployment verification requirements and coordinate releases across Cloud Platform, Backend, Autonomy, and Frontend teams.
  • Maintain audit-ready deployment and release records.
  • Own the incident management framework, including declarations, severity levels, escalation paths, incident command, communications, and notifications.
  • Run incident command during significant events and lead blameless post-incident reviews.
  • Manage government-sponsor notification obligations for incidents affecting authorized systems.
  • Distinguish incidents from problems and drive analysis of recurring issues.
  • Maintain authoritative records of deployed systems, versions, digests, dependencies, baselines, and configuration drift.
  • Produce audit-ready evidence for the ISSO and track POA&M remediation commitments.
  • Translate security and compliance controls into practical engineering processes.
  • Build automation to reduce manual compliance work and improve operational evidence reliability.
  • Establish effective change, release, incident, reliability, compliance, and remediation processes within the first 12 months.

Requirements:

  • 8+ years of relevant experience across change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related disciplines.
  • Demonstrated experience leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and performance management.
  • Strong understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.
  • Experience establishing and operating change, release, and incident management processes in complex technical environments.
  • Proven ability to coordinate complex initiatives across engineering teams and stakeholders you do not directly manage.
  • Technical fluency sufficient to evaluate and challenge engineering impact assessments, deployment strategies, rollback plans, and root-cause analyses.
  • Ability to remain calm, decisive, and directive during active incidents and make sound decisions under pressure.
  • Strong written communication skills, particularly for incident communications, post-incident reports, operational documentation, and executive updates.
  • Ability to create scalable processes that provide appropriate control without unnecessarily slowing engineering teams.
  • Strong ownership, judgment, and comfort operating in a fast-moving and ambiguous environment.
  • Must be a U.S. Citizen and able to obtain and maintain a U.S. Government security clearance.
  • Prior incident command experience in a regulated, defense, government, or safety-relevant environment.
  • Experience operating cloud systems subject to U.S. Government authorization or compliance requirements.
  • Familiarity with POA&Ms, security authorization processes, configuration baselines, and audit evidence management.
  • Experience implementing GitOps-based change and release processes.
  • Knowledge of ITIL practices or equivalent hands-on experience developing effective service management processes.
  • Experience with tools such as Jira, PagerDuty, status pages, runbook platforms, and incident management systems.
  • Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.

Benefits:

  • 100% Employer paid Health, Dental and Vision Insurance for you and your families
  • Life Insurance (Employer Paid)
  • Ability to participate in the companies 401k program (Matching)
  • Unlimited PTO policy with an enforced 2 week minimum
  • Equity Package
  • Work / Home Office Stipend
  • Global Entry
  • 16 Week Paid Parental Leave
  • Monthly Health and Wellness Stipend
  • Bonus