Staff DevOps Engineer
Posted 1hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Staff DevOps Engineer owning AWS, EKS, IaC, CI/CD, observability, security, and reliability for MEDvidi’s AI-powered mental healthcare platform. Leading hands-on infrastructure strategy in a fully remote EU collaboration model.
Responsibilities:
- Set and drive the technical vision and quarterly roadmap for infrastructure, with clear trade-offs and measurable goals
- Run and evolve AWS and Kubernetes (EKS) infrastructure, including cluster management, autoscaling, policy enforcement, and zero-downtime operations
- Own Infrastructure as Code using Terraform and AWS CDK in TypeScript
- Own GitLab CI/CD, including reusable/shared templates, OIDC, and self-managed GitLab
- Build and own practical observability using Prometheus, Grafana, OpenTelemetry, OpenSearch, and CloudWatch
- Maintain log pipelines and APM as a shared view of system health
- Operate blue-green deployments with health-gated automated rollback
- Own zero-downtime PostgreSQL schema migrations and CI migration gating
- Lead security engineering in a HIPAA environment, including secrets hygiene, PHI-aware logging and data handling, leak scanning, and Vault managed as code
- Partner directly with product teams to remove infrastructure friction and improve developer experience
- Use agentic AI as a core workflow, integrating autonomous-agent output into production
- Set technical direction, ship infrastructure changes hands-on, and own reliability, cost, security, performance, deployment health, and developer velocity outcomes
Requirements:
- 6+ years of experience in DevOps/infrastructure engineering
- Strong systems fundamentals, Linux administration, and troubleshooting, including performance analysis, resource management, and process debugging
- Hands-on AWS experience with EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, and S3
- Production Kubernetes/EKS experience, including cluster management, node scaling, and policy enforcement; Karpenter, Kyverno, or similar
- Strong Infrastructure as Code experience with Terraform and AWS CDK in TypeScript
- CI/CD ownership with GitLab CI/CD, reusable/shared templates, OIDC id_tokens, and self-managed GitLab
- Practical monitoring and observability experience with Prometheus, Grafana, OpenTelemetry, OpenSearch, and CloudWatch; log shipping and error tracking/APM
- Practical security engineering experience with secrets rotation, short-lived credentials, leak scanning, and PHI-aware logging
- HashiCorp Vault as code experience, including KV, JWT/OIDC authentication for CI, and policy design
- Experience with blue-green deployments and automated, health-gated rollback
- PostgreSQL zero-downtime schema migration experience using expand/contract and migration gating in CI
- Containers experience with Docker, ECR, immutable tags, and image lifecycle
- Network and protocol fundamentals, including load balancing, TLS, and DNS
- Hands-on agentic AI workflows, such as Claude Code or similar, including delegating to autonomous agents and integrating their output into production
- Developer-focused mindset and strong problem-solving for complex system issues
- Strong technical writing, including docs-as-code, ADRs, and design documents via MRs
- Fluent Russian and English (B2)
- Experience working effectively in remote, distributed teams
- Preferred: experience in HIPAA, SOC 2, or similar regulated environments; Ansible; Node.js operations; GitOps tooling; deeper PostgreSQL administration; AWS certifications
Benefits:
- Competitive compensation package
- Fully remote long-term collaboration under a B2B model
- Health insurance after the probation period
- Sports & wellness compensation
- Personalized English lessons via Preply
- 19 paid vacation days annually
- 4 additional wellness days each year
- Paid sick leave for the first 5 working days
- Thoughtful gifts for key life events
- Offline corporate events
















