Senior Monitoring Analyst

Posted 13hrs ago

Employment Information

Industry
Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Analista NOC sênior monitorando redes críticas de telecomunicações para provedores e operadoras. Liderando incidentes, troubleshooting e evolução de ambientes Zabbix e Grafana.

Responsibilities:

  • Monitor availability, performance and capacity of assets, links and critical services
  • Lead analysis of alarms, events, metrics and indicators in high-impact or complex incidents
  • Perform diagnostics and technical escalations, identifying causes, impacts, priorities and containment actions
  • Execute advanced troubleshooting of connectivity, infrastructure, collection and monitoring services
  • Design and maintain hosts, templates, items, preprocessing, triggers, macros, discovery rules (LLD), actions and event correlation in Zabbix
  • Create and evolve dashboards, variables, datasources, transformations and alerts in Grafana
  • Investigate collection failures via SNMP, ICMP, agents, Syslog, APIs, proxies and other integrations
  • Define and review thresholds, dependencies, severities and suppression rules to increase alarm accuracy
  • Analyze capacity and availability trends, anticipating risks, saturation and degradations
  • Conduct root cause analyses and propose corrective and preventive actions
  • Ensure logging, escalations and technical communications according to defined SLAs and workflows
  • Mentor analysts, perform technical reviews and share knowledge
  • Define standards, document solutions and contribute to the evolution of monitoring and NOC operations

Requirements:

  • Advanced experience with Zabbix in production environments
  • Strong knowledge of distributed architecture, proxies, templates, items, preprocessing, LLD (low-level discovery), triggers, dependencies, macros, actions and event correlation
  • Advanced experience with Grafana, including datasources, variables, transformations, alerts and dashboards
  • Knowledge of SNMPv2c/v3, ICMP, agents, Syslog, APIs and webhooks
  • Solid understanding of TCP/IP, addressing, DNS, VLANs, switching, routing, latency and packet loss
  • Ability to interpret metrics, establish baselines and correlate events and trends
  • Proficiency with tools such as ping, traceroute, MTR, dig, curl and tcpdump
  • Linux knowledge for analyzing services, logs, processes, resources and integrations
  • Experience with incident and problem management, root cause analysis, escalation, SLAs and operating at scale
  • Strong background in monitoring, infrastructure or NOC operations, working on critical incidents and high-availability environments
  • Hands-on mastery of Zabbix and Grafana, including configuration, troubleshooting and evolution of production environments

Benefits:

  • Continuous development
  • Technical support
  • Incentives for training and certifications relevant to the role