Senior DevOps Engineer – AI Cloud
Posted 1ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior DevOps Engineer building and operating AI cloud infrastructure for Bitdeer’s Bitcoin mining and AI computing platforms. Automating deployments, Kubernetes infrastructure, GPU workloads, reliability, and security.
Responsibilities:
- Design, implement, and maintain end-to-end CI/CD pipelines for software applications and machine learning models
- Automate build, testing, deployment, and rollback processes
- Build, optimize, and scale cloud-native infrastructure using Kubernetes and Docker
- Manage and provision specialized computing resources, including GPU clusters, for AI workloads and model inferencing
- Own high-availability design in production environments
- Implement disaster recovery, self-healing mechanisms, capacity planning, and performance tuning
- Champion Infrastructure as Code practices using Terraform, Ansible, and Helm
- Architect and refine monitoring, logging, and alerting systems
- Collaborate with R&D, Data Science, Security, and Business teams on workflow optimization and Platform Engineering initiatives
- Establish and enforce system stability and security standards, release workflows, Zero Trust access controls, secrets management, and compliance
- Lead troubleshooting during complex anomalies and major incidents, conduct root cause analysis, and implement preventative remediation plans
Requirements:
- Bachelor's degree or above in Computer Science, Engineering, or a related technical field
- 5+ years of hands-on experience in DevOps, Site Reliability Engineering, or Cloud Infrastructure roles
- Expert-level knowledge of Linux operating systems and core networking principles, including TCP/IP, DNS, HTTP, load balancing, and VPCs
- Deep mastery of Docker and Kubernetes orchestration, cluster management, and production best practices
- Proficiency designing and managing infrastructure on major public or hybrid cloud platforms, including AWS, GCP, Azure, or Alibaba Cloud
- Experience with multi-cloud and hybrid-cloud strategies
- Strong coding and scripting capabilities in at least one major language, such as Go, Python, or Shell
- Practical understanding of CI/CD, Infrastructure as Code, observability, and Site Reliability Engineering principles
- Exceptional problem-solving abilities, technical judgment, and cross-team communication skills
- a• Preferred: familiarity with MLOps, model serving/inferencing frameworks, GPU clusters, large-scale distributed systems, Internal Developer Platforms, Zero Trust, DevSecOps, SOC2, ISO27001, and technical leadership


















