NOC Engineer / SRE
Posted 4ds ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
SRE-NOC Engineer ensuring 24/7 service reliability for NICE’s AI and cloud software products. Automating operations, improving observability, and responding to production incidents.
Responsibilities:
- Act as a primary or escalation responder in a 24x7 on-call rotation
- Lead or support Major Incident response, including triage, mitigation, and resolution
- Coordinate across Engineering, Infrastructure, Security, and Product teams
- Execute and improve runbooks, playbooks, and escalation paths
- Drive blameless post-incident reviews and track corrective actions
- Own service health monitoring across infrastructure, applications, and dependencies
- Design and maintain alerting strategies aligned with SLIs/SLOs
- Reduce alert fatigue through signal-to-noise improvements
- Build dashboards using Grafana, Prometheus, Datadog, Splunk, and/or CloudWatch
- Automate repetitive operational tasks and reduce manual toil
- Improve mean time to detect and mean time to resolve
- Develop scripts and tools in Python, Bash, Go, or similar
- Implement self-healing and auto-remediation where possible
- Partner with engineering teams to improve system reliability
- Support and troubleshoot Linux systems, cloud platforms, and Kubernetes/containerized environments
- Assist with capacity planning and availability reviews
- Ensure operational readiness for production releases
Requirements:
- Strong Linux systems administration
- Experience with incident management and production support
- Familiarity with cloud infrastructure, preferably AWS
- Familiarity with Docker and Kubernetes
- Familiarity with monitoring and alerting platforms
- Scripting or programming experience in Python, Bash, Go, or similar
- Understanding of networking fundamentals, including DNS, TCP/IP, and load balancing
- Experience working in 24x7 NOC or production operations environments
- Ability to handle high-pressure incidents calmly and effectively
- Strong written and verbal communication for incident coordination
- Comfort working from runbooks and improving them when needed
- Experience defining or operating to SLOs/SLIs preferred
- Prior migration from traditional NOC to SRE model preferred
- Infrastructure as Code experience with Terraform, Ansible, or similar preferred
- Exposure to security, compliance, or regulated environments preferred















