Principal HPC Network Engineer
Posted 10hrs ago
Employment Information
Report this job
Job expired or something wrong with this job?
Job Description
Senior HPC Networking Engineer responsible for high-performance networking infrastructure design and troubleshooting. Working with InfiniBand technologies and Fortinet solutions in AI-driven workloads.
Responsibilities:
- Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.
- Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.
- Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.
- Perform performance tuning, monitoring, and capacity planning for HPC networking systems.
- Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).
- Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.
- Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.
- Develop and maintain documentation for network architecture, configurations, and operational procedures.
- Participate in on-call rotations and provide escalation support for critical incidents.
- Lead or contribute to network upgrades, migrations, and new deployments.
Requirements:
- 5+ years of experience in network engineering, with a focus on HPC or data center environments.
- Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).
- Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.
- Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).
- Experience with network performance analysis and troubleshooting tools.
- Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).
- Strong analytical and problem-solving skills.
- Preferred: Experience with large-scale HPC clusters or AI/ML infrastructure.
- Knowledge of RDMA, MPI, and low-latency networking concepts.
- Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.
- Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).
Benefits:
- Operate some of the most advanced AI infrastructure environments in production today.
- Work with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.
- Help define operational standards and reliability practices for next-generation AI infrastructure services.
- Influence the adoption of AI-powered operational capabilities through k0rdent AI.
- Work alongside highly skilled engineers solving complex infrastructure and platform challenges at scale.
- Join a growing organisation investing heavily in AI infrastructure, platform services, and operational innovation.

















