Senior RDMA Performance Engineer

Posted 13hrs ago

Employment Information

Education
Salary
Experience
Job Type

Report this job

Job expired or something wrong with this job?

Job Description

RDMA Performance Engineer optimizing NVIDIA AI data center networks for CodiLime, a software and network engineering services company. Benchmarking RDMA fabrics and tuning congestion control for high-performance workloads.

Responsibilities:

  • Configure, troubleshoot, and maintain NVIDIA/Mellanox Spectrum switches and ConnectX Network Interface Cards (NICs)
  • Implement and tune RoCEv2, Priority-based Flow Control (PFC), and Explicit Congestion Notification (ECN) for AI workloads
  • Optimize network environments for maximum throughput, minimal latency, and minimal Job Completion Time (JCT)
  • Conduct performance benchmarking and validation testing on AI-ready data center fabrics using RDMA and NVIDIA Collective Communication Library (NCCL)
  • Use AI agents and specialized tools to automate routine network tasks, perform testing, and analyze network telemetry
  • Author and maintain technical documentation covering network designs, configurations, and benchmarking results
  • Work as part of an Agile project team in a startup environment on a US-client project

Requirements:

  • 7+ years of professional experience in network engineering
  • Preferably 3+ years specifically focused on data center network design and architecture
  • Hands-on experience configuring and troubleshooting BGP, EVPN/VXLAN, LACP, and ECMP
  • Deep understanding of AI-dedicated data center designs, including Rail-Optimized Design (ROD) and Rail-Unified Design (RUD) topologies, with practical hands-on configuration and deployment experience
  • Proven experience configuring, managing, and troubleshooting NVIDIA/Mellanox Spectrum switches and ConnectX Network Interface Cards (NICs)
  • Deep understanding of RoCEv2, Priority-based Flow Control (PFC), and Explicit Congestion Notification (ECN)
  • Experience tuning switch buffers and implementing Quality of Service (QoS) policies
  • Experience conducting performance benchmarking and validation for AI-ready data center fabrics using RDMA and NVIDIA Collective Communication Library (NCCL)
  • Ability to leverage AI agents/tools to automate tasks, testing, and analysis
  • Strong technical documentation skills and professional English communication
  • Professional-level certifications such as CCNP/CCIE Enterprise/DC, JNCIP, or equivalent are nice-to-haves
  • Hands-on experience with pyATS, NAPALM, Batfish, Nornir, or NetBox is a nice-to-have
  • Exposure to GitLab CI, GitHub Actions, Jenkins, and infrastructure-as-code concepts is a nice-to-have

Benefits:

  • Flexible working hours and approach to work: fully remotely, in the office or hybrid
  • Professional growth supported by internal training sessions and a training budget
  • Solid onboarding with a hands-on approach to give you an easy start
  • A great atmosphere among professionals who are passionate about their work
  • The ability to change the project you work on