Senior DevOps Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior DevOps Engineer operating Kubernetes, cloud, and GPU inference infrastructure for an award-winning AI product company. Building secure, observable CI/CD platforms across cloud and on-premises environments.
Responsibilities:
- Deploy, operate, and evolve a microservices-based platform running in Kubernetes clusters across AWS, GCP, and on-premises Rancher
- Operate and support GPU-based ML inference services using Triton Inference Server and vLLM on RunPod, Scaleway, and Nebius
- Build and maintain Docker images for all microservices and ensure a stable service lifecycle
- Maintain and scale development and production Kubernetes clusters
- Participate in deployment debugging, incident investigation, and performance troubleshooting
- Develop, maintain, and evolve custom Helm charts for each service
- Design and operate CI/CD pipelines using GitHub and GitLab for on-premises customer deployments
- Ensure platform compliance with SOC 2 requirements and improve security and compliance processes
- Manage cluster access via NetBird VPN and implement role-based access control using group policies
- Deploy and manage infrastructure using Terraform and Ansible
- Develop and continuously improve observability systems using Grafana, Prometheus, and the ELK stack
- Continuously optimize infrastructure across IaC, IAM, observability, and CI/CD
- Work with Python, Kubernetes, Linux, Docker, GitHub CI/CD, PostgreSQL, ClickHouse, Kafka, Superset, Terraform, and Ansible
Requirements:
- Minimum 5 years of experience in a DevOps and/or Site Reliability Engineering role
- Strong hands-on experience with Linux system administration
- Extensive experience deploying, operating, and scaling Kubernetes in both cloud and bare-metal environments
- Deep expertise and practical experience with at least one major cloud provider, preferably Google Cloud Platform
- Experience with ML inference on GPU/CPU is a strong plus
- Proven experience implementing SRE practices and building observability stacks using Grafana, Prometheus, and Loki
- Strong adherence to GitOps, Infrastructure as Code (IaC), and CI/CD principles
- Advanced expertise in Terraform, Ansible, and Python
- Ability to work in high-uncertainty environments and rapidly learn new technologies and patterns
- Proactive mindset and ability to debug and understand the product beyond DevOps tasks
- Strategic thinking for selecting technologies and architectural approaches based on long-term goals
Benefits:
- Fully remote
- 21 vacation days + public holidays + 5 sick days
- Private English lessons via Preply
- Fast career progression
- Startup pace with enterprise stability — real clients, real revenue, no bureaucracy
- Cutting-edge tech stack
- High engineering bar and real ownership
- Opportunity to work on award-winning AI products
- Exposure to Speech Technologies, NLP, Generative AI, LLMs, diffusion models, and voice-first agentic architecture


















