Lead Site Reliability Engineer
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Lead Site Reliability Engineer modernizing cloud, Kubernetes, and deployment infrastructure for Intellum’s corporate education technology platform. Driving reliability, observability, incident response, and Systems Engineering mentorship.
Responsibilities:
- Own and drive infrastructure modernization initiatives, including evolution from legacy compute environments to container-orchestrated infrastructure.
- Design and maintain infrastructure as code across multiple cloud providers.
- Improve CI/CD systems and deployment tooling for efficient, observable, and recoverable releases.
- Provide technical leadership through mentorship, architecture guidance, knowledge sharing, and engineering-practice support.
- Establish and evolve SLI and SLO practices, monitoring, alerting, and load-testing capabilities.
- Lead platform incident response, troubleshooting, root cause analysis, and corrective actions.
- Drive visibility into cloud infrastructure costs and incorporate cost considerations into architecture decisions.
- Improve developer experience through infrastructure, development environments, deployment workflows, and production feedback loops.
- Partner with Security and Engineering on access controls, infrastructure hardening, compliance, and secure infrastructure practices.
- Identify operational and infrastructure risks, recommend priorities, and help drive the Systems Engineering technical roadmap.
- Mentor engineers and contribute to developing the Systems Engineering team and technical practices.
- Perform other duties as assigned.
Requirements:
- 8+ years of hands-on experience in infrastructure, DevOps, platform engineering, site reliability engineering, or a related discipline, including experience building and operating production systems.
- Deep hands-on experience designing, operating, and troubleshooting highly available production infrastructure.
- Production experience across more than one major cloud provider, with depth in at least one of AWS or Google Cloud and working fluency in the other.
- Significant experience with container orchestration and Kubernetes in production environments, including cluster operations, workload configuration, reliability, and troubleshooting.
- Experience modernizing production infrastructure, including migrations from VM-based or legacy environments toward containerized or cloud-native architectures.
- Strong infrastructure-as-code experience using Terraform or comparable tooling, with an emphasis on repeatability and automation.
- Experience building, operating, or significantly improving CI/CD systems and deployment infrastructure.
- Strong incident response and troubleshooting capabilities, including experience diagnosing complex distributed-system failures and contributing to effective post-incident review.
- Strong Linux administration skills and scripting or programming ability in Ruby, Python, or a comparable language.
- Experience working in a SaaS environment where reliability, availability, and production stability are critical.
- Ability to collaborate effectively with distributed teams across US and European time zones and participate in an on-call rotation.
- Strong communication skills and the ability to provide technical direction, mentor other engineers, and influence infrastructure decisions across teams.
- Bachelor's degree in a related field or equivalent practical experience; equivalent experience is genuinely accepted for this role.
- Production experience across both AWS and Google Cloud simultaneously is preferred.
- Prior leadership or management, cloud cost management or FinOps, Spinnaker/Jenkins, Ruby on Rails, SOC 2, AI-assisted development tooling, or learning technology experience is preferred.
Benefits:
- Medical - 100% of employee premiums for selected individual plans
- Dental - 100% of employee premiums covered
- Vision - 100% of employee premiums covered
- LinkedIn Learning
- 401(k) plus matching (US Based Only)
- Flexible PTO
- Calm subscription
- Annual Company Retreat
- Personal development budgets












