Site Reliability Engineer – SRE

Posted 4hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Site Reliability Engineer operating Kubernetes-based identity security services for Software Mind’s enterprise clients. Monitoring production, resolving incidents, improving CI/CD, automation, resilience, and observability.

Responsibilities:

  • Support the deployment, operation, and ongoing maintenance of the Karuna service running on Kubernetes
  • Monitor production environments to ensure high availability, reliability, and performance
  • Investigate, troubleshoot, and resolve production incidents, performing root cause analysis where appropriate
  • Analyze application logs and debug production issues using Splunk
  • Perform first-level troubleshooting of UI-related issues involving Web Components, collaborating with front-end engineers when deeper investigation is required
  • Support deployment activities and contribute to maintaining and improving CI/CD pipelines
  • Identify opportunities for automation and operational improvements
  • Work closely with engineering teams and technical stakeholders in an international environment to improve operational processes, enhance service resilience, and optimize observability
  • Take ownership of operational tasks and proactively drive issues to resolution while working independently with minimal supervision

Requirements:

  • Commercial experience as a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or in a similar role
  • Hands-on experience supporting production services running on Kubernetes
  • Experience monitoring distributed applications and responding to production incidents
  • Practical knowledge of log analysis and troubleshooting using Splunk or similar monitoring tools
  • Understanding of cloud-native applications and modern operational practices
  • Familiarity with CI/CD pipelines and deployment processes
  • Basic understanding of Web Components and the ability to perform first-level UI troubleshooting
  • Strong analytical and problem-solving skills
  • Ability to work independently, prioritize tasks, and make sound technical decisions in ambiguous situations
  • Excellent communication skills and confidence collaborating with distributed engineering teams
  • Very good spoken and written English
  • Experience with one or more major cloud platforms
  • Experience supporting enterprise SaaS or security-focused products

Benefits:

  • Flexible employment and remote work
  • International projects with leading global clients
  • International business trips
  • Non-corporate atmosphere
  • Language classes
  • Internal & external training
  • Private healthcare and insurance
  • Multisport card
  • Well-being initiatives