Openstack Infrastructure Engineer
Posted 9hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
OpenStack Infrastructure Engineer building and scaling xneelo’s reliable web-hosting compute and storage infrastructure. Operating OpenStack, Ceph, Linux, networking, and infrastructure automation systems.
Responsibilities:
- Build, manage, and scale OpenStack compute and storage infrastructure
- Operate and troubleshoot open-source infrastructure systems
- Design and maintain reliable compute and storage hosting platforms
- Troubleshoot compute, storage, physical and virtual networking, databases, message queues, containers, hypervisors, and guest workloads
- Plan and execute upgrades, migrations, and high-impact infrastructure changes
- Automate infrastructure using Infrastructure as Code, configuration management, and version-controlled tooling
- Manage architecture, access control, operational workflows, capacity, performance, recovery, and data-loss exposure
- Produce architecture decisions, runbooks, change plans, incident findings, and recovery procedures
- Participate in technical reviews and contribute to innovation and complex IT infrastructure problem-solving
- Participate in on-call support and respond to occasional out-of-hours incidents, maintenance, or operational requests
Requirements:
- Experience operating open-source infrastructure systems; extensive experience expected at senior level
- OpenStack implementation, operations, or troubleshooting experience
- Experience with Ceph or similar distributed storage systems, including capacity, replication, failure domains, recovery, latency, and performance
- Strong Linux systems administration and troubleshooting skills
- Practical knowledge of data-centre-grade hardware, including enterprise servers, CPUs, memory, storage media, NICs, firmware, and out-of-band management
- Understanding of how hardware design, power, cooling, rack layout, and component failure affect platform reliability and performance
- Strong networking fundamentals; experience in OVN, OVS, BGP underlays, LACP, Juniper, IPv4, IPv6, and physical or virtual cloud networks advantageous
- Ability to troubleshoot across compute, storage, physical and virtual networking, databases, message queues, containers, hypervisors, and guest workloads
- Experience with virtualisation and cloud infrastructure at scale
- Security-minded approach to architecture, automation, access control, and operational workflows
- SRE mindset focused on reducing failure probability, recovery time, operational risk, and data-loss exposure
- Experience with Infrastructure as Code, configuration management, and automation practices
- Preference for repeatable, version-controlled automation over undocumented manual work
- Discipline planning and executing upgrades, migrations, and high-impact infrastructure changes
- Strong end-to-end ownership from physical infrastructure and network fabric through OpenStack services to customer workloads
- Ability to reason about capacity and performance across CPU, memory, storage, IOPS, latency, throughput, packet rates, and control-plane scale
- Ability to use AI-assisted tools productively while understanding their limitations
- Clear technical documentation skills, including architecture decisions, runbooks, change plans, incident findings, and recovery procedures
- Ability to contribute constructively to technical reviews, challenge unsafe assumptions, and respond well to detailed feedback
- Ability to work effectively in a distributed, largely asynchronous team
- Willingness to participate in an on-call rotation managed through Rootly and respond to occasional out-of-hours incidents, maintenance, or operational requests
- Experience with a reasonable number of Ubuntu Linux, Ansible, OpenStack, Ceph, Terraform, Rundeck, LXC, Nspawn, OVN, OVS, BGP, JunOS, LACP, Docker, Prometheus, Grafana, Git, Jira, or similar tools
Benefits:
- Salary negotiable, commensurate with skills and track record
- High level of discretion and autonomy guided by company principles and values
- Distributed, largely asynchronous team
- Participation in an on-call rotation managed through Rootly



















