Senior/Staff DevOps Engineer

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior/Staff DevOps Engineer owning AWS, Kubernetes, security, and observability for MEDvidi’s AI mental healthcare platform. Driving reliable infrastructure and developer experience remotely from Portugal.

Responsibilities:

  • Set and drive the technical vision and quarterly roadmap for infrastructure with clear trade-offs and measurable goals
  • Run and evolve AWS and Kubernetes (EKS) infrastructure, including cluster management, autoscaling with Karpenter, policy enforcement with Kyverno, and zero-downtime operations
  • Own Infrastructure as Code end to end using Terraform and AWS CDK in TypeScript
  • Own GitLab CI/CD, including reusable/shared templates, OIDC, and self-managed GitLab
  • Build and own practical observability using Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch, log pipelines, and APM
  • Maintain fast blue-green deployments and health-gated automated rollback
  • Own zero-downtime PostgreSQL schema migrations using expand/contract and CI migration gating
  • Own security engineering in a HIPAA environment, including secrets hygiene, credential rotation, short-lived credentials, leak scanning, PHI-aware log and data handling, and Vault managed as code
  • Partner directly with product teams to remove infrastructure friction and improve developer experience
  • Use agentic AI as a core part of the workflow and integrate autonomous-agent output into production
  • Set technical direction, ship infrastructure hands-on, and own outcomes for reliability, cost, security, performance, deployment health, and developer experience

Requirements:

  • 6+ years in DevOps/infrastructure engineering
  • Strong systems fundamentals
  • Solid Linux administration and troubleshooting, including performance analysis, resource management, and process debugging
  • Hands-on AWS experience with EC2, EKS, RDS, ElastiCache, Lambda, SQS, EventBridge, API Gateway, ALB, and S3
  • Production Kubernetes/EKS experience, including cluster management, node scaling, and policy enforcement; Karpenter, Kyverno, or similar
  • Strong Infrastructure as Code experience with Terraform and AWS CDK in TypeScript
  • CI/CD ownership with GitLab CI/CD, including reusable/shared templates, OIDC id_tokens, and self-managed GitLab
  • Practical monitoring and observability experience with Prometheus, Grafana, OpenTelemetry, OpenSearch, CloudWatch, log-shipping, and error tracking/APM
  • Practical security engineering experience with secrets rotation, short-lived credentials, leak scanning, and PHI-aware logging
  • HashiCorp Vault as code experience, including KV, JWT/OIDC authentication for CI, and policy design
  • Blue-green deployments with automated, health-gated rollback
  • PostgreSQL zero-downtime schema migrations using expand/contract and migration gating in CI
  • Containers experience with Docker, ECR, immutable tags, and image lifecycle
  • Network/protocol fundamentals including load balancing, TLS, and DNS
  • Hands-on agentic AI workflows, such as Claude Code or similar
  • Developer-focused mindset and strong problem-solving for complex system issues
  • Strong technical writing, including docs-as-code, ADRs, and design docs via MRs
  • Fluent Russian and English (B1)
  • Experience working effectively in remote, distributed teams
  • Preferred: experience in regulated/compliance-heavy environments such as HIPAA or SOC 2
  • Preferred: Ansible for VM fleet management
  • Preferred: Node.js application operations with pm2 and npm
  • Preferred: GitOps tooling such as ArgoCD or Flux and deeper PostgreSQL database administration
  • Preferred: AWS certifications

Benefits:

  • Competitive compensation package
  • Health insurance after the probation period
  • Sports & wellness compensation
  • Personalized English lessons via Preply
  • 19 paid vacation days annually
  • 4 additional wellness days each year
  • Paid sick leave for the first 5 working days
  • Thoughtful gifts for key life events
  • Offline corporate events
  • Fully remote long-term collaboration under a B2B model