NOC Engineer / SRE

Posted 4ds ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

SRE-NOC Engineer ensuring 24/7 service reliability for NICE’s AI and cloud software products. Automating operations, improving observability, and responding to production incidents.

Responsibilities:

  • Act as a primary or escalation responder in a 24x7 on-call rotation
  • Lead or support Major Incident response, including triage, mitigation, and resolution
  • Coordinate across Engineering, Infrastructure, Security, and Product teams
  • Execute and improve runbooks, playbooks, and escalation paths
  • Drive blameless post-incident reviews and track corrective actions
  • Own service health monitoring across infrastructure, applications, and dependencies
  • Design and maintain alerting strategies aligned with SLIs/SLOs
  • Reduce alert fatigue through signal-to-noise improvements
  • Build dashboards using Grafana, Prometheus, Datadog, Splunk, and/or CloudWatch
  • Automate repetitive operational tasks and reduce manual toil
  • Improve mean time to detect and mean time to resolve
  • Develop scripts and tools in Python, Bash, Go, or similar
  • Implement self-healing and auto-remediation where possible
  • Partner with engineering teams to improve system reliability
  • Support and troubleshoot Linux systems, cloud platforms, and Kubernetes/containerized environments
  • Assist with capacity planning and availability reviews
  • Ensure operational readiness for production releases

Requirements:

  • Strong Linux systems administration
  • Experience with incident management and production support
  • Familiarity with cloud infrastructure, preferably AWS
  • Familiarity with Docker and Kubernetes
  • Familiarity with monitoring and alerting platforms
  • Scripting or programming experience in Python, Bash, Go, or similar
  • Understanding of networking fundamentals, including DNS, TCP/IP, and load balancing
  • Experience working in 24x7 NOC or production operations environments
  • Ability to handle high-pressure incidents calmly and effectively
  • Strong written and verbal communication for incident coordination
  • Comfort working from runbooks and improving them when needed
  • Experience defining or operating to SLOs/SLIs preferred
  • Prior migration from traditional NOC to SRE model preferred
  • Infrastructure as Code experience with Terraform, Ansible, or similar preferred
  • Exposure to security, compliance, or regulated environments preferred