Global Resolution Engineer
Posted 2hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Global Resolution Engineer resolving complex networking issues for WEKA’s AI-native data infrastructure. Driving HPC, AI/ML, and enterprise deployment performance, scalability, and support readiness.
Responsibilities:
- Serve as the primary escalation point for complex networking, fabric design, connectivity, and performance issues in WEKA deployments
- Troubleshoot Layer 1–4 networking problems in high-throughput, low-latency AI/ML and HPC environments
- Mentor Customer Success Engineers, review case strategy, and elevate team networking competence
- Analyze field data to identify emerging patterns, bottlenecks, and systemic network configuration flaws
- Maintain and evolve WEKA’s Known Issues Database, Technical Advisories, and Knowledge Base
- Partner with Development on root cause analyses and future product capabilities and diagnostic tools
- Conduct advanced networking training for internal support teams and key customers
- Participate in technical design reviews with internal and customer architecture teams
- Improve internal tooling for packet capture, latency analysis, flow visualization, and path validation
- Represent GRE in cross-functional architecture reviews focused on networking readiness and scalability
- Participate in global 24/7 on-call operations and resolve high-severity incidents
Requirements:
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related technical field; equivalent experience accepted
- 15+ years in advanced technical support, network engineering, or network architecture roles in enterprise, cloud, or HPC environments
- Demonstrated ability to act as a technical lead in multi-team escalations and customer-critical situations
- Experience working with or supporting AWS, Azure, or GCP and hybrid deployments preferred
- English fluency required; additional languages a plus
- Flexibility to support global operations, including on-call rotation and occasional travel
- Subject matter expertise in Ethernet, Infiniband, RDMA/RoCE, DPDK, and UCX
- Extensive experience with Arista, NVIDIA/Mellanox, Cisco, and Juniper switching and routing platforms
- Deep understanding of TCP/IP, LACP, ECMP, jumbo frames, MTU tuning, congestion control, and QoS
- Hands-on experience with ethtool, iperf, perfquery, ifconfig, and fabric utilities
- Advanced Linux systems administration, including sysctl, iptables/nftables, and routing tables
- Proven ability to resolve complex technical issues in distributed systems and storage-heavy, compute-intensive workloads
- Ability to lead architectural discussions on datacenter fabrics for scale-out systems
- Strong documentation skills and ability to produce internal guides and customer-facing technical content
- Bonus experience with Kubernetes networking, Terraform, Ansible, and Python/Bash automation
Benefits:
- Equal opportunity employer and nondiscrimination policy
- Inclusive and authentic workplace commitment
- Global 24/7 operations with on-call participation
- Occasional travel









