Senior Site Reliability Engineer – Network Observability

Posted 20hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

Senior SRE operating enterprise network observability platforms for Blueprint Technologies, a cloud, AI, and data solutions firm. Automating operations, troubleshooting telemetry systems, and improving reliability at scale.

Responsibilities:

  • Support the operation and reliability of enterprise network observability platforms that ingest and analyze network telemetry
  • Administer and operate Linux and Windows virtual machines hosted in a cloud environment
  • Maintain compliance with security, configuration, and operational standards
  • Operate and support network observability platforms, primarily Syslog-NG and Trapd
  • Execute operating system, application, and security upgrades and patching
  • Investigate automated alerts, customer-reported incidents, and platform performance issues
  • Troubleshoot network observability configurations, software applications, and operating systems
  • Create and maintain processing rules using regular expressions
  • Support enterprise network performance and application monitoring platforms
  • Analyze traffic patterns, network telemetry, resource utilization, system performance, and security events with network and security engineers
  • Conduct system capacity planning and recommend scalability and reliability improvements
  • Automate recurring operational tasks using Bash, PowerShell, or Python
  • Deploy, configure, and manage cloud services for reliability, scalability, security, and cost efficiency
  • Implement DevOps practices including CI/CD pipelines, source control, and infrastructure as code
  • Perform business continuity and disaster recovery failover testing
  • Manage assigned projects and program components according to objectives and timelines
  • Participate in daily stand-ups and collaborate with engineering teams
  • Participate in an on-call rotation and provide incident response
  • Deliver compliance, incident-resolution, and project-delivery outcomes

Requirements:

  • Bachelor’s degree in computer science, computer engineering, information technology, or a related technical field, or equivalent professional experience
  • Five to seven years of enterprise experience in systems engineering, network engineering, site reliability engineering, or a related IT infrastructure role
  • At least five years of Linux system administration experience
  • At least three years of network engineering experience
  • At least three years of hands-on Syslog-NG experience
  • Strong understanding of SNMP, SNMP Traps, NetFlow, and gNMI
  • Strong knowledge of enterprise networking, including routing and switching protocols
  • Hands-on experience with Azure or a comparable cloud platform
  • Experience operating and troubleshooting network monitoring systems in a large enterprise environment
  • Experience with system capacity planning, configuration management, compliance audits, and performance analysis
  • Ability to investigate complex incidents involving infrastructure, operating systems, applications, and network telemetry
  • Strong project management, collaboration, and communication skills
  • Preferred: proficiency in Bash, PowerShell, or Python for automation
  • Preferred: strong proficiency with regular expressions
  • Preferred: experience with IBM SevOne Network Performance Manager, Broadcom AppNeta, or comparable observability platforms
  • Preferred: experience using source-control platforms and development workflows
  • Preferred: working knowledge of Ansible and Ansible playbooks
  • Preferred: intermediate knowledge of KQL, T-SQL, or comparable data-retrieval languages
  • Preferred: experience implementing CI/CD pipelines and infrastructure-as-code practices
  • Preferred: exposure to commercially available artificial intelligence platforms
  • Preferred: experience connecting network telemetry with AI-enabled workflows

Benefits:

  • Medical, dental, and vision coverage
  • Flexible Spending Account (FSA)
  • 401(k) retirement plan
  • Competitive paid time off
  • Parental leave
  • Professional growth and development opportunities