darrenqu.net

AI Networking

RoCE Congestion Control on ConnectX-5: Capturing a Real CNP, and Why the Sender Can't Fake CE

2377 words 12 min read

roce

2,000 CE-marked packets in, 2,000 CNPs out. A byte-level look at a real ConnectX-5 CNP, the NIC rewriting ECN behind my back, and the DCQCN roles that soft-RoCE was missing.

on this page

This picks up where RoCE Congestion Control Without RDMA NICs left off. There, soft-RoCE (rxe) took DCQCN apart layer by layer. Everything up to and including CE-mark delivery reproduced in software. The last two layers did not: 3,682 CE-marked packets went in and zero CNPs came back, because rxe implements the RoCE data transport but not the congestion-control paths.

That left two open questions:

  • On a real NIC, does the receiver (the NP) actually answer CE with a CNP?
  • Is the CNP’s BTH opcode 0x81 or 0x80? The specification, the IETF drafts, NVIDIA’s documentation, DPDK and Scapy all say 0x81; Wireshark’s opcode table labels 0x80 as CNP. Last time I could only read sources, not measure.

A friend lent me two servers with ConnectX-5 NICs for about an hour. The plan was to set CE at the sender, the way I had on rxe, and let the receiver do the rest. The NIC stopped that at the first step.

Lab scope. Two RHEL 8.8 hosts (kernel 4.18.0-477), MLNX_OFED 23.10-3.2.2.0, single-port 100G ConnectX-5, back to back with no switch. NIC firmware version was not recorded. Both hosts are production Kubernetes nodes; only a dedicated 100G NIC was used and the cluster’s own interfaces were not touched. Everything below describes this setup. Other NIC models, firmware and driver versions may behave differently.

The setup#

Node1 (sender)                                Node2 (receiver)
┌─────────────────────┐                     ┌─────────────────────┐
│ mlx5_4 / ens9np0    │═══ 100G, direct ════│ mlx5_2 / ens9np0    │
│ 10.0.0.1            │                     │ 10.0.0.2            │
└─────────────────────┘                     └─────────────────────┘
      RP role                                       NP role

ib_write_bw writes from client to server, so Node1 runs the client and plays the RP — it sends data, receives CNPs and is supposed to slow down. Node2 runs the server and plays the NP — it receives CE-marked packets and sends CNPs. That decides where each counter has to be read: np_* counters on Node2, rp_* counters on Node1.

DCQCN was enabled on all eight priorities on both hosts, with identical parameters. The MLNX_OFED sysfs tree, trimmed to the entries this article uses:

roce_np/cnp_dscp:48
roce_np/cnp_802p_prio:6
roce_np/min_time_between_cnps:4
roce_np/enable/0:1
roce_rp/enable/0:1

min_time_between_cnps deserves a note before any traffic runs. It rate-limits CNP generation (in microseconds), and only an implementation that really sends CNPs needs such a knob. rxe has nothing corresponding to it.

The counters make the same point. rxe exposes 18, none related to ECN, CE or CNP. ConnectX-5’s hw_counters has 30, and four of them exist for DCQCN alone: np_ecn_marked_roce_packets, np_cnp_sent, rp_cnp_handled and rp_cnp_ignored.

Verify the capture first#

Same rule as last time: prove the instrument before trusting its readings. RoCE bypasses the kernel, so tcpdump on the netdev sees none of the RDMA traffic. libpcap 1.9 and later can capture directly from an RDMA device instead:

30.mlx5_0 (RDMA sniffer)
32.mlx5_1 (RDMA sniffer)
34.mlx5_4 (RDMA sniffer)

One run of traffic through tcpdump -i mlx5_4 caught UDP 4791 with 0 packets dropped by kernel. The capture path works.

It also nearly misled me. The first 20 packets of a run contained no data packets at all:

IP 10.0.0.1.55402 > 10.0.0.2.4791: UDP, length 280
IP 10.0.0.2.55402 > 10.0.0.1.4791: UDP, length 280
IP 10.0.0.1.55402 > 10.0.0.2.4791: UDP, length 280
IP 10.0.0.1.55402 > 10.0.0.2.4791: UDP, length 32
IP 10.0.0.2.55402 > 10.0.0.1.4791: UDP, length 20
...

The 280-byte packets are the rdma_cm exchange; the 20-byte ones are ACKs. The rxe lab had already shown that the CM exchange happens before the QP exists, so the TOS option never applies to it and it always reads tos 0x0. Judging ECN from these packets would have concluded “the NIC clears ECN entirely”, which is wrong. From here on, every TOS check filters on udp[4:2] > 1000 to look at data packets only.

One more constraint worth knowing: an RDMA device supports one sniffer at a time. Starting a second tcpdump on the same device while one is running fails with Failed to create flow for device mlx5_4.

Setting CE at the sender: the NIC rewrites it#

On rxe, rdma_cm passes user-supplied ECN bits through untouched, so --tos=3 puts CE (ECN 11) on the wire. With no switch available to mark packets, the obvious move was to do the same here.

Node1 sent with --tos=3. Node2 received tos 0x2 — ECT(0):

IP (tos 0x2,ECT(0), ttl 64, id 19931, offset 0, flags [DF], proto UDP (17), length 1084)
    10.0.0.1.54573 > 10.0.0.2.4791: UDP, length 1056
IP (tos 0x2,ECT(0), ttl 64, id 19932, offset 0, flags [DF], proto UDP (17), length 1068)
    10.0.0.1.54573 > 10.0.0.2.4791: UDP, length 1040

Two guesses followed. Perhaps the whole TOS byte was being rewritten, so I tried a value with a DSCP in it. Perhaps the NIC forces ECT(0) only because the RP is enabled, so I temporarily disabled the RP on priority 0 of Node1 and tried again:

--tosConditionTOS received on Node2DSCPECN
3RP enabled0x2010 ECT(0)
0x83RP enabled0x8232, preserved10 ECT(0)
3RP disabled0x2010 ECT(0)

In this ConnectX-5 + MLNX_OFED 23.10-3.2.2.0 setup, the NIC was observed to rewrite the two ECN bits of RoCE data packets to ECT(0), leave the DSCP alone, and do so regardless of whether the RP is enabled.

The rewrite happens on the sending side rather than the receiver clearing CE. The next experiment indicates as much: Node2’s np_ecn_marked_roce_packets does count CE-marked packets when they reach it by another route.

The RP was re-enabled immediately after that run. Every later test ran with the RP on.

Getting CE past the sender’s QP path#

What the NIC rewrites is traffic leaving through the RDMA QP send path. So what happens if a CE-marked RoCE packet never goes through a QP, and instead leaves through the kernel network stack like any other Ethernet frame?

The procedure:

  1. Keep an ib_write_lat test running between the two hosts, so a live QP exists.
  2. On Node1, capture one real data packet that Node1 itself sent, using the RDMA sniffer.
  3. Set the ECN bits in its IP header to 11 and recompute the IP header checksum.
  4. Transmit it 2,000 times from ens9np0 through a raw socket, about 1 ms apart.

The captured packet was a complete RDMA WRITE Only:

FieldValue
BTH opcode0x0a, RC RDMA WRITE Only
UDP length1064 = 8 + 12 (BTH) + 16 (RETH) + 1024 + 4 (ICRC)
IP TOS0x2, ECT(0) — already rewritten by the NIC
DestQP0x35, the QP on Node2

One precondition has to hold: the packet must still be valid after the ECN change. The RoCEv2 ICRC covers the IP header, but the TOS byte is masked to all ones when the ICRC is computed, so changing ECN leaves the ICRC valid. The IP header checksum does need recomputing. The UDP checksum is zero in RoCEv2 and can be ignored.

The replayed packet carries a stale PSN, and the receiver’s transport layer should treat it as a duplicate. So this was a bet on ordering: does the NP look at CE before or after the PSN check? If after, all 2,000 packets would be discarded as duplicates and not one CNP would come back.

2,000 CE in, 2,000 CNPs out#

Counters after the run. Node1 has a saved pre-injection baseline, with every counter below at zero. Node2 has no pre-injection snapshot; that gap is covered in its own section below.

CounterNode1 (RP)Node2 (NP)
np_ecn_marked_roce_packets2000 ⚠️2000
np_cnp_sent02000
rp_cnp_handled20000
rp_cnp_ignored00
out_of_sequence / packet_seq_err0 / 00 / 0

Put the counters next to the two captures and the evidence lines up:

  • Node2’s NP recognized 2,000 CE-marked packets and sent 2,000 CNPs in response to replayed packets. At the very least, in this setup, CNP generation does not require the packet to pass the normal PSN receive check first. Whether ConnectX-5 evaluates ECN before the PSN check or in an independent path is something this experiment cannot tell apart.
  • Node1’s rp_cnp_handled rose by 2,000 with rp_cnp_ignored at zero, so every one of those CNPs entered the RP’s processing path and was accepted. That proves the RP received and processed them. It does not, on its own, prove the send rate went down.
  • Each capture holds 2,000 CNPs: Node2 on transmit, Node1 on receive.

One detail closes the loop. The injected packets were addressed to Node2’s QP 0x35. All 2,000 CNPs came back with DestQP 0x33, which is Node1’s QP (local address: QPN 0x0033 in the ib_write_lat output). The NP built CNPs aimed at the peer QP from its existing QP context, rather than echoing back the DestQP of the packet that triggered them.

The CNPs are spaced 1.056 ms apart at the median, matching the injection pace. Every CE-marked packet bought exactly one CNP. min_time_between_cnps is 4 µs, far below the 1 ms injection interval, so the rate limit never engaged and nothing was coalesced.

Anatomy of a real CNP#

All 2,000 CNPs are field-for-field identical. One of them, raw:

0x0000:  45c2 003c ece6 4000 4011 3906 0a00 0002
0x0010:  0a00 0001 0000 12b7 0028 0000 8100 ffff
0x0020:  4000 0033 0000 0000 0000 0000 0000 0000
0x0030:  0000 0000 0000 0000 8bc6 384b

Decoded:

FieldBytesMeaning
IP TOSc2DSCP 48 + ECT(0). DSCP 48 matches cnp_dscp:48 in sysfs
UDP source port00000
UDP dest port / length / checksum12b7 / 0028 / 00004791; 40 = 8 + 32; no checksum
BTH opcode81CNP
BTH P_Keyffffdefault partition
BTH bytes 4–74000 0033BECN = 1, DestQP = 0x33
BTH PSN0000 00000
Reserved16 zero bytesthe specification requires the sender to zero them
ICRC8bc6 384ba real ICRC

A few results deserve their own paragraph.

0x81 or 0x80 is now settled. All 2,000 CNPs carry 0x81; not one carries 0x80. That agrees with the specification, the IETF drafts, NVIDIA’s documentation, DPDK and Scapy. The outlier is Wireshark: version 4.2.2, used here, has no name for opcode 0x81 and decodes the CNP as Opcode: Unknown (129).

Wireshark 4.2.2 decoding a captured CNP: Opcode Unknown (129), Reserved 40, and the reserved bytes and ICRC shown as Unknown/Vendor Specific Data

Wireshark 4.2.2 on one of the 2,000 captured CNPs.

Look at the line Reserved: 40 as well. That byte is not reserved: its top two bits are FECN and BECN, and 0x40 is BECN = 1 — the bit that makes this packet a congestion notification. It is the second Wireshark blind spot the previous article described from the dissector source, now on a real CNP: the only congestion signal in the header is displayed as a reserved field.

The UDP source port is 0. Ordinary RoCEv2 data packets use the source port to carry per-flow entropy for ECMP hashing — this capture’s data packets used ports such as 54573 and 59399. With the CNP’s source port fixed at 0, a network hashing on the usual five-tuple gets much less transport-layer entropy from CNPs than from data packets, and CNPs between the same pair of endpoints may concentrate on one path. This lab is a direct link, so that was not tested on a real multipath fabric.

The CNP itself is ECN-capable. Its TOS is 0xc2, ECT(0), which means a switch on the return path could mark the CNP CE as well.

The structure matches what the rxe article derived. A 32-byte UDP payload — BTH 12 + reserved 16 + ICRC 4 — with BECN set. Last time the layout came from the specification by hand; now there is a real packet to compare it against.

Side by side with rxe#

LayerWhatrxe (previous article)ConnectX-5 (this article)
L1–L2RDMA works, RoCEv2 on the wire✅✅
L3DSCP settable✅✅
L4ECN-capable✅ set by the user✅ forced by the NIC
L5CE marking✅ by tc RED⚠️ not possible at the sender; replayed through the kernel
L6CNP returned❌ zero✅ 2,000 → 2,000
L7RP receives the CNP—✅ rp_cnp_handled = 2000
L7RP cuts the rate—❌ not tested

The previous article’s reading holds: ConnectX-5 implements the DCQCN NP and RP processing paths, and rxe implements neither.

Where this stops#

The L7 rate cut was not tested. This article shows the RP received and processed CNPs, not that the send rate actually dropped. Testing that needs a real congestion point producing CE continuously — a switch with ECN marking, or a BlueField DPU sitting in the forwarding path. A direct link cannot provide one.

Node1’s np_ecn_marked_roce_packets also reads 2,000. Node1 is the sender, so this counter should only have risen on Node2. “The CNPs themselves carry CE” is ruled out, since the CNP TOS is 0xc2, ECT(0). One possibility is that the NIC counted the 2,000 CE-marked frames on the transmit path as they left Node1. That has not been verified.

Node2 has no pre-injection baseline. Its counter snapshot from before the injection was not saved. Three things support attributing all 2,000 to this run: no earlier test managed to put CE on the wire; Node2’s counters equal the injected count exactly; and both captures hold exactly 2,000 CNPs. Strictly speaking, though, it is a gap in the evidence chain.

Conclusions#

In this back-to-back ConnectX-5 setup:

  • The NIC was observed to rewrite the ECN field of RoCE data packets to ECT(0) — only the two ECN bits, not the DSCP, and independently of the RP setting. Faking CE at the sender works on rxe and does not work here.
  • The NP is real: it answers CE-marked packets with CNPs one for one, and builds them for the peer QP from its existing QP context.
  • The RP received and processed every CNP: rp_cnp_handled is 2,000 and rp_cnp_ignored is 0. The rate cut itself was not tested.
  • The CNP opcode is 0x81, with BECN set, a 32-byte UDP payload, DSCP taken from cnp_dscp, and UDP source port 0.

Put simply: in these experiments rxe reproduces the path up to CE and the CP’s behavior, while ConnectX-5 has complete NP and RP processing paths. ConnectX-5’s normal RoCE send path also forces ECN to ECT(0), so CE cannot be faked at the sender the way it can on rxe.


References

Observed on ConnectX-5 with MLNX_OFED 23.10-3.2.2.0 on RHEL 8.8, captures decoded with Wireshark 4.2.2. Firmware version not recorded. Other NICs, firmware and driver versions may behave differently.

← more in AI Networking