Senior Site Reliability Engineer – Kubernetes

Posted 5hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior SRE securing Kubernetes reliability for Software Mind’s ServiceNow AI platform. Managing observability, incident response, deployments, and Node.js/JVM production troubleshooting.

Responsibilities:

  • Support the deployment, operation, and reliability of production services running on Kubernetes
  • Monitor service health and investigate production incidents across distributed applications
  • Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements
  • Troubleshoot application runtime, networking, and service-to-service issues with engineering teams
  • Support CI/CD, GitOps-based deployments, observability, and production monitoring
  • Work within a client-directed backlog and established priorities
  • Own production reliability for the AI Experience Framework stack end to end, including Kubernetes, observability, and troubleshooting Node.js and JVM systems

Requirements:

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering, or a closely related role
  • Strong recent hands-on experience supporting Kubernetes-based production services
  • 3+ years of hands-on production Kubernetes experience strongly preferred
  • Kubernetes production operations, including deployment, scaling, rollout/rollback, resource tuning, and service-to-service troubleshooting
  • Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene
  • Splunk experience for log aggregation, search, and production troubleshooting
  • Prometheus and Grafana experience building alert rules and dashboards
  • CI/CD and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux
  • Strong Linux and networking fundamentals, including DNS, load balancing, TCP/HTTP, HTTP/2, and Kubernetes networking
  • Production troubleshooting across Node.js and JVM/Java services, with strong depth in at least one runtime environment
  • Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based authentication
  • Very good spoken and written English
  • Web Components/Lit experience, server-side rendering or isomorphic runtime experience, canary rollout/multi-version production operations, distributed tracing, KEDA or event-driven autoscaling, and enterprise platform integration experience are additional skills

Benefits:

  • Flexible employment and remote work
  • International projects with leading global clients
  • International business trips
  • Non-corporate atmosphere
  • Language classes
  • Internal & external training
  • Private healthcare and insurance
  • Multisport card
  • Well-being initiatives