Senior Ceph/Rook-Ceph, Kubernetes Storage Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Storage Engineer operating production Ceph and Kubernetes storage infrastructure for a long-term project. Troubleshooting clusters, performance, failures, and high-availability environments.
Responsibilities:
- Join a long-term project as a Senior Ceph/Rook-Ceph + Kubernetes Storage Engineer
- Operate and troubleshoot production Ceph environments
- Work with Kubernetes storage and infrastructure troubleshooting
Requirements:
- 5+ years in DevOps, SRE, infrastructure or storage engineering
- Deep hands-on production Ceph experience, including OSD, MON, MGR, PGs, recovery, backfill, capacity planning, performance tuning, scaling and high availability
- Experience building, operating, upgrading and troubleshooting production Ceph clusters
- Hands-on Rook-Ceph in Kubernetes, preferably in business-critical production environments
- Strong Kubernetes knowledge, including CSI/storage integration
- Linux administration and troubleshooting at OS/hardware level
- Automation experience with Ansible
- Experience with Kubernetes tooling such as Helm
- Monitoring/observability with Prometheus and Grafana
- Experience investigating storage performance, latency, disk failures, network bottlenecks and recovery issues
- Strong understanding of failure domains, CRUSH topology, replication and storage architecture
- Good English communication skills
- OpenStack experience, especially Ceph integration with Cinder, Glance, Nova and RBD (highly valuable but not mandatory)
- Petabyte-scale Ceph environments (highly valuable but not mandatory)
- Bare-metal infrastructure (highly valuable but not mandatory)
- Large production Rook-Ceph clusters (highly valuable but not mandatory)
- CephFS and RGW/S3 experience (highly valuable but not mandatory)
- Customer-facing troubleshooting/support experience (highly valuable but not mandatory)


















