RoCE Congestion Control Without RDMA NICs: Soft-RoCE Gives You One of DCQCN's Three Roles
tc marks 3,682 packets CE. All 3,682 arrive at the receiver. Soft-RoCE sends back zero CNPs. Where the line falls between what you can verify in software and what needs hardware — and how I nearly got the answer wrong.
on this page
Every explanation of DCQCN assumes you have RDMA NICs. Every soft-RoCE guide stops at “ib_write_bw returns a number.” I could not find anyone who joins the two up, so the question goes unanswered: without RDMA hardware, how much of RoCE congestion control can you actually verify?
DCQCN has three roles#
The feedback loop needs three participants:
| Role | Who plays it | What it does |
|---|---|---|
| CP — Congestion Point | the switch | marks the IP header ECN field as CE when the queue builds |
| NP — Notification Point | the receiving NIC | sees CE, sends a CNP back to the sender |
| RP — Reaction Point | the sending NIC | receives the CNP, cuts the rate for that QP |
The CP runs on a switch, but the NP and RP run inside NIC hardware. DCQCN’s rate control is a per-QP state machine with its own timers and recovery curve, implemented in silicon — not something a driver computes in software.
The lab, and which parts of it are honest#
One VM. One veth pair, one end moved into a network namespace. rdma_rxe on both ends, tc in the middle.
netns test1 init_net (host)
┌────────────────┐ ┌────────────────┐
│ veth-a rxe0 │══════════════════│ veth-b rxe1 │
│ 172.16.0.1/24 │ tbf + red ecn │ 172.16.0.2/24 │
└────────────────┘ └────────────────┘
NP role CP role RP roleBefore running anything, define what each piece represents:
| Component | Stands in for | Faithful? |
|---|---|---|
rdma_rxe (soft-RoCE) | an RDMA NIC | data transport yes, hardware no. The RoCEv2 data and ACK wire formats are exercised here, but there is no kernel bypass, no CNP generation and no DCQCN state machine |
| veth pair | a cable | no real buffer, no DCB, no PFC |
tc red ... ecn | a switch’s WRED + ECN marking | faithful for the marking behavior tested here. Linux RED genuinely rewrites the ECN field on queue build-up, but it is not a physical switch buffering model |
The ECN-marking function playing the congestion point is real. The parts playing the NICs are not. That determines what the lab can prove and the order of the tests below.
Lab scope. Ubuntu 26.04, kernel 7.2.6-070206-generic, rdma-core 61.0-2ubuntu3, single VM, 4 vCPU / 2 GB. Soft-RoCE only — no RDMA hardware was involved, and nothing here should be extrapolated to real NIC behavior.
Building it#
Everything below runs on a single VM. No RDMA hardware, no second machine. The one read-only checks in the first block need no privileges; everything after them does.
Check the kernel version first. It is the cheapest gate and it decides whether any of the
rest is possible — on 7.0 the topology below builds without error and then quietly does the
wrong thing, putting both RDMA devices in init_net:
uname -r # needs 7.1 or newer
grep RDMA_RXE /boot/config-$(uname -r) # expect CONFIG_RDMA_RXE=mUbuntu 26.04 ships 7.0, which is not enough; everything here ran on a mainline build,
7.2.6-070206-generic. The reason 7.1 is the floor is at the end of the article.
sudo modprobe rdma_rxe
sudo ip netns add test1
sudo ip link add veth-a type veth peer name veth-b
sudo ip link set veth-a netns test1
sudo ip netns exec test1 ip addr add 172.16.0.1/24 dev veth-a
sudo ip netns exec test1 ip link set veth-a up
sudo ip netns exec test1 ip link set lo up
sudo ip addr add 172.16.0.2/24 dev veth-b
sudo ip link set veth-b up
sudo ip netns exec test1 rdma link add rxe0 type rxe netdev veth-a
sudo rdma link add rxe1 type rxe netdev veth-bNow confirm the namespace really got its own RDMA stack. Both sides must show a listener
on 4791: that per-namespace UDP socket is precisely what Linux 7.1 added, so this is the
runtime proof, where uname -r was only the claim. A vendor kernel can report a new
version without carrying the patch.
sudo ip netns exec test1 ss -Huln 'sport == :4791' # must not be empty
sudo ss -Huln 'sport == :4791'
sudo ip netns exec test1 rdma link # must show 'netdev veth-a'
sudo ip netns exec test1 ping -c 2 -W 1 172.16.0.2If the namespace side is empty, or rdma link there shows nothing while the host shows
both devices, stop — the kernel is too old and the rest will not mean anything.
Then generate traffic. -R uses rdma_cm for connection setup, which selects the RoCEv2
GID automatically; --tos=2 sets DSCP 0 with ECT(0), which is what makes RED mark
rather than drop:
sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1Start every run by killing the previous one. A server left waiting from a failed attempt will happily answer the next client that comes along, and you will be measuring something other than what you think — more on that at the end.
sudo pkill -f ib_write_bw; sudo pkill -f tcpdumpThat is all you need between runs. The pkill -f above matches every ib_write_bw and
tcpdump process on the VM, so use it only on a dedicated lab VM; on a shared host, record
the process IDs and stop only what this experiment started.
Everything from here to the end of the article assumes the topology is still up, so leave it in place until you are finished. The block below removes it:
sudo pkill -f ib_write_bw; sudo pkill -f tcpdump
sudo tc qdisc del dev veth-b root
sudo ip netns exec test1 rdma link del rxe0
sudo rdma link del rxe1
sudo ip link del veth-b
sudo ip netns del test1A reboot has the same effect: namespaces, veth pairs and rxe devices are all gone afterwards. Re-run the build block above and carry on.
Verify the instrument before the subject#
The rule I followed, learned the hard way on an earlier EVPN lab: check that the measurement works before trusting what it says. Each layer is a precondition for the next.
L1 RDMA works → L2 it really is RoCEv2 on the wire → L3 DSCP is settable → L4 packets are ECN-capable → L5 CE marks get produced → L6 CNPs come back → L7 the rate drops.
Layer 3 is easy to skip. Setting a DSCP value and reading it back is not a test of RoCE; it verifies the capture path before later results depend on it.
L1–L5: everything up to and including the switch works#
L1. ib_write_bw between the two devices: 527.37 MiB/sec. Soft-RoCE is CPU-bound — at full speed exactly three threads saturate, the client, the server, and the rxe kworker.
More cores only move this number; the L5 and L6 runs are shaped to 100 Mbit and barely
register, so a smaller VM reaches the same conclusions with a different L1 figure.
#bytes #iterations BW peak[MiB/sec] BW average[MiB/sec] MsgRate[Mpps]
65536 5000 527.37 485.77 0.007772L2. Capture the exchange and look at what is actually on the wire:
mkdir -p ~/roce-lab/evidence/
sudo tcpdump -i veth-b -nn -U -w ~/roce-lab/evidence/L2-roce-v2.pcap udp port 4791 &
sleep 1
sudo ip netns exec test1 ib_write_bw -d rxe0 -R >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R -n 5000 172.16.0.1 | tail -3
sudo pkill -INT tcpdump261,403 packets, every one of them on UDP 4791. Three dominant length clusters account for 261,385 of them:
| Count | UDP length | Breakdown | Type |
|---|---|---|---|
| 253,314 | 1040 | 12 (BTH) + 1024 + 4 (ICRC) | RDMA Write Middle/Last |
| 4,025 | 1056 | 12 (BTH) + 16 (RETH) + 1024 + 4 | RDMA Write First |
| 4,046 | 20 | 12 (BTH) + 4 (AETH) + 4 (ICRC) | Mostly Acknowledge |
The header sizes are fixed by the specification, and only one assignment makes the bytes add up — RETH carries the remote address, rkey and length, so it can only ride on a First. That is arithmetic, though, not measurement. Counting opcodes gives a second, independent answer:
for op in 0x06 0x07 0x08 0x11; do
printf 'opcode %-5s : %8s\n' "$op" \
"$(tcpdump -r ~/roce-lab/evidence/L2-roce-v2.pcap -nn "udp[8] == $op" 2>/dev/null | wc -l)"
doneopcode 0x06 : 4025 RDMA Write First
opcode 0x07 : 249292 RDMA Write Middle
opcode 0x08 : 4022 RDMA Write Last
opcode 0x11 : 4040 AcknowledgeThe data-packet counts line up to the packet. 4,025 Firsts is exactly the 1056-byte cluster. 249,292 Middles plus 4,022 Lasts is 253,314 — exactly the 1040-byte cluster. The 20-byte cluster is mostly Acknowledges: 4,040 of 4,046 carry opcode 0x11. Six carry a different opcode, and another 18 packets fall outside the three length clusters above. I did not chase those 24 remaining packets further.
65536 ÷ 1024 bytes of RDMA payload = 64 packets per message, and the opcodes show how those 64 divide: one First carrying the RETH, then 62 Middles, then one Last. 249,292 ÷ 4,025 = 61.94, and First, Last and Acknowledge agree with one another to within half a percent — one of each per message. The residual differences are small enough to be capture boundaries rather than protocol behavior.
L3 and L4. perftest’s -T/--tos sets the TOS byte via rdma_cm. The open question was whether the low two bits survive — many stacks mask the ECN field out of a user-supplied TOS, on the grounds that ECN belongs to the transport. Run it four times, changing only
--tos, and read the byte back off the wire each time:
for T in "" 104 2 106; do
sudo pkill -f ib_write_bw >/dev/null 2>&1; sudo pkill -f tcpdump >/dev/null 2>&1; sleep 1
F=~/roce-lab/evidence/tos-${T:-unset}.pcap
sudo tcpdump -i veth-b -nn -U -w "$F" udp port 4791 & sleep 1
sudo ip netns exec test1 ib_write_bw -d rxe0 -R ${T:+--tos=$T} >/dev/null 2>&1 & sleep 2
sudo ib_write_bw -d rxe1 -R ${T:+--tos=$T} -n 2000 172.16.0.1 >/dev/null 2>&1
sudo pkill -INT tcpdump; sleep 1
printf '=== --tos=%s ===\n' "${T:-unset}"
tcpdump -r "$F" -nn -v 2>/dev/null | grep -oE 'tos 0x[0-9a-f]+' | sort | uniq -c
doneThey do not get masked:
--tos | observed on the wire | packets | DSCP | ECN |
|---|---|---|---|---|
| unset | 0x0 | 28,599 | 0 | 00 Not-ECT |
| 104 | 0x68 | 32,190 | 26 | 00 |
| 2 | 0x2 | 31,501 | 0 | 10 ECT(0) |
| 106 | 0x6a | 30,816 | 26 | 10 ECT(0) |
The packet counts differ between runs only because a broad capture at full rate drops most of what it sees — 75 to 78 percent in these four. The TOS value is the measurement; the volume is not.
RED drops Not-ECT packets rather than marking them, so without ECT there is no L5. Because rdma_cm passes the ECN bits through, no packet forging is needed.
One structural detail: in each of the three runs where the data carries a
non-zero TOS, exactly 39 packets still read tos 0x0 — the same 39 every time. They are
consistent with the rdma_cm exchange, which happens before the QP exists and therefore before
the TOS option can apply to it. In the unset run there is no separate 39, because the
exchange and the data are both 0x0 and cannot be told apart. Confirming the opcodes and
sequence of those 39 packets would turn that inference into direct evidence.
L5. Shape well below the ~485 MiB/s the unshaped run managed — call it 4 Gbit/s — so that a queue actually forms:
sudo tc qdisc add dev veth-b root handle 1: tbf rate 100mbit burst 32kb latency 50ms
sudo tc qdisc add dev veth-b parent 1:1 handle 10: red \
limit 400000 min 30000 max 100000 avpkt 1000 burst 55 ecn bandwidth 100mbitOmitting bandwidth is a quiet trap — RED prints RED: set bandwidth to 10Mbit and computes its averaging constants against the wrong link rate.
Now run traffic through it and read the qdisc counters. Unlike a packet capture, they are not affected by tcpdump drops:
sudo pkill -f ib_write_bw; sudo pkill -f tcpdump; sleep 1
sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1 | tail -3
sudo tc -s qdisc show dev veth-bShaped to 100 Mbit, 5000 iterations of 64 KB take about half a minute instead of the second they took unshaped. The queue must build before RED has anything to act on.
qdisc tbf 1: root refcnt 129 rate 100Mbit burst 32Kb lat 50ms
Sent 346324396 bytes 320040 pkt (dropped 0, overlimits 834591 requeues 0)
qdisc red 10: parent 1:1 limit 400000b min 30000b max 100000b ecn
Sent 346324396 bytes 320040 pkt (dropped 0, overlimits 3698 requeues 0)
marked 3698 early 0 pdrop 0 other 0Read all three numbers together. marked means CE marks were produced. early 0 pdrop 0 means nothing was dropped — proof that ECT(0) took effect, because RED recognized the packets as ECN-capable and marked them instead.
marked is a counter on the sending qdisc. It proves marks were applied, not that they survived the veth traversal and reached the receiver. So measure inside the namespace, on the ingress side:
sudo pkill -f ib_write_bw; sudo pkill -f tcpdump; sleep 1
sudo tc -s qdisc show dev veth-b | grep marked # baseline, counters are cumulative
sudo ip netns exec test1 tcpdump -i veth-a -nn -s 128 -B 65536 \
-w ~/roce-lab/evidence/ce-arrival.pcap 'udp port 4791 and (ip[1] & 0x03) == 3' &
sleep 1
sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1 | tail -3
sudo pkill -INT tcpdump; sleep 1
sudo tc -s qdisc show dev veth-b | grep marked # subtract the baseline
tcpdump -r ~/roce-lab/evidence/ce-arrival.pcap -nn | wc -lThe capture runs inside the namespace on veth-a, the receiving side of the link. The filter
(ip[1] & 0x03) == 3 selects packets whose IP header carries ECN 11 — Congestion
Experienced.
| Measurement point | Value |
|---|---|
qdisc marked, before this run | 3,698 |
qdisc marked, after | 7,380 |
| delta — marks applied by the sender | 3,682 |
| CE-marked packets received inside the namespace | 3,682 |
An exact match, and the capture that produced the second figure dropped nothing:
3682 packets captured / 3682 received by filter / 0 dropped by kernel. Every mark RED
applied on the way out arrived at the far end still carrying it. The congestion signal
demonstrably reached rxe0.
The CP’s ECN-marking function is reproducible in software. Linux tc acts on the same IP
ECN field, while its queue and buffer model remains different from a physical switch ASIC.
L6: 3,682 CE marks in, zero CNPs out#
Capturing the reverse direction needs a narrow filter and a small snaplen, or the kernel
starts dropping — which for a rare packet would be fatal. udp[8] is the first byte of
the UDP payload, which is the BTH opcode:
sudo pkill -f ib_write_bw; sudo pkill -f tcpdump; sleep 1
rm -f ~/roce-lab/evidence/L6-reverse.pcap # never census a capture from an older run
sudo tcpdump -i veth-b -nn -s 128 -B 65536 \
-w ~/roce-lab/evidence/L6-reverse.pcap 'udp port 4791 and src 172.16.0.1' &
sleep 1
sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1 | tail -3
sudo pkill -INT tcpdump; sleep 1
tcpdump -r ~/roce-lab/evidence/L6-reverse.pcap -nn | wc -l
for i in $(seq 0 255); do
op=$(printf '0x%02x' $i)
c=$(tcpdump -r ~/roce-lab/evidence/L6-reverse.pcap -nn "udp[8] == $op" 2>/dev/null | wc -l)
[ "$c" -gt 0 ] && printf 'opcode %-5s : %s\n' "$op" "$c"
doneAnd back the other way. 5,036 packets matched the filter and 5,020 reached the file — the sixteen in between were still in tcpdump’s buffer when it was signaled, and the kernel dropped none at all. All 256 opcode values were then swept across those 5,020:
| BTH opcode | Meaning | Count |
|---|---|---|
0x11 | Acknowledge | 5,009 |
0x04 | RC SEND Only (perftest sync) | 9 |
0x64 | UD SEND Only (rdma_cm control) | 2 |
0x80 / 0x81 | CNP | 0 |
5,020 of 5,020 accounted for, with nothing left over.
CNP is opcode 0x81. Wireshark’s InfiniBand dissector disagrees: its bth_opcode_tbl maps 0x80 to CNP and has no entry at 0x81, with the opcode field declared unmasked. A real CNP therefore appears as an unnamed opcode 129, and filtering on the display name finds nothing. The specification (0b10000001), the IETF drafts, NVIDIA’s documentation (decimal 129, “RoCEv2 Congestion Management Ack”), DPDK’s BTH flow-matching code and Scapy’s RoCE module all say 0x81. Scapy provides a quick check — python3 -c "from scapy.contrib.roce import CNP_OPCODE; print(hex(CNP_OPCODE))":
CNP_OPCODE = 0x81
def cnp(dqpn):
return BTH(opcode=CNP_OPCODE, becn=1, dqpn=dqpn) / CNPPadding()DPDK’s patch is explicit about exactly this use case:
flow create 0 group 1 ingress pattern eth / ipv4 / udp dst is 4791 /
ib_bth opcode is 0x81 dst_qp is 0xd3 / end actions queue index 0 / endI swept both values regardless. Zero at each.
Wireshark has a second blind spot in the same area, and this one bites anyone doing congestion work on real hardware too. Here is a real RDMA WRITE Middle packet from this lab, dissected in full:
InfiniBand
Base Transport Header
Opcode: Reliable Connection (RC) - RDMA WRITE Middle (7)
0... .... = Solicited Event: False
.0.. .... = MigReq: False
..00 .... = Pad Count: 0
.... 0000 = Header Version: 0
Partition Key: 65535
Reserved: 00
Destination Queue Pair: 0x000022
0... .... = Acknowledge Request: False
.000 0000 = Reserved (7 bits): 0
Packet Sequence Number: 16516294Notice how much of that is broken out to individual bits. Solicited Event, MigReq, Pad
Count and Header Version each get their own bitmask line; so does Acknowledge Request, and
even the seven reserved bits beside it. Then look at the line that just says Reserved: 00.
That byte is not reserved. Bits 7 and 6 are FECN and BECN, InfiniBand’s own congestion
notification bits — the second of the two congestion mechanisms RoCE carries. The kernel
knows they are there; rxe_hdr.h defines BTH_FECN_MASK (0x80000000) and
BTH_BECN_MASK (0x40000000) against exactly this word. Wireshark does not decode them at
all: the strings fecn and becn appear nowhere in its InfiniBand dissector, and the byte
is added to the tree as an undifferentiated Reserved. There is no field to filter on.
To see those two bits you have to read the byte yourself — udp[12] & 0x80 for FECN,
udp[12] & 0x40 for BECN. In this capture the byte reads 00, which is what the driver
source predicts: those accessors exist and nothing ever calls them.
The receiver’s counters provide another check. rdma statistic show link rxe0/1 exposes eighteen counters — packets, bytes, sequence errors, RNR errors, retries — and not one of them relates to ECN, CE or CNP. rxe does not merely fail to respond to congestion marks; it does not count them.
And the driver source says why. rxe_resp.c, the responder path where an NP role would live, never reads the IP header ECN bits and never generates a CNP. The InfiniBand BTH carries its own congestion bits, FECN and BECN, and rxe does define accessors for them in rxe_hdr.h — bth_fecn(), bth_becn(), bth_set_fecn(), bth_set_becn(). Checking the 37 source files in drivers/infiniband/sw/rxe/, those accessors are referenced zero times outside their own definitions. They are inherited header plumbing that nothing calls.
The tested kernel and the mainline 7.3-rc source snapshot checked for this article both lack these paths.
Where the line falls#
Verifiable in software: the RoCEv2 data and ACK wire formats exercised here, DSCP marking, ECT(0) transmission, and CE-mark generation and delivery. L1 through L5, with a complete evidence chain.
Not verifiable: the other two roles of the DCQCN loop — the receiver returning a CNP, and the sender reacting to it.
Why: those roles are implemented in NIC hardware. In kernel 7.2.6 and the mainline 7.3-rc snapshot examined here, Soft-RoCE implements the data transport but not the DCQCN congestion-control paths.
The boundary is not simply hardware versus software; it is between IP-layer marking and RoCE transport behavior.
ECN lives in the IP header — two bits, defined by RFC 3168.
Marking CE is an IP-layer action, which is why Linux tc can reproduce the marking behavior
tested here while acting on a field in a layer it already owns.
CNP lives one layer up, in the RoCE transport. Generating one means reading a field in the IP header and then emitting a new RoCE transport packet carrying a particular BTH opcode. That crossing — IP signal in, RoCE transport packet out — is what the examined rxe versions do not implement. The IP-header marking path tested here reproduces in software. In the RoCE transport, soft-RoCE gives you the data format but not the congestion-control behavior.
PFC is further out of reach still. veth has no DCB and no priority queues, so 802.1Qbb cannot be approached at all.
A richer simulator would have to model the NIC behavior. NVIDIA’s DSX Air
markets “emulated hardware fidelity to the network adapter”
and “simulation of NVIDIA ConnectX SuperNICs”. On a standard Ubuntu host node there,
lspci returns QEMU devices and nothing else.
00:00.0 Host bridge: Intel Corporation 440FX - 82441FX PMC [Natoma]
00:08.0 Ethernet controller: Red Hat, Inc. Virtio network device [1af4:1000]
...
driver: virtio_netNo Mellanox device, no mlx5. RoCE on that node uses soft-RoCE again, with the same missing
paths. I did not survey every node type DSX Air offers, and the documentation has no page
describing what is and is not simulated, so treat this as one data point rather than a verdict
on the product.
Without hardware, you can still practise RoCE protocol structure, QoS marking and switch-side ECN configuration on one small VM. The DCQCN feedback loop is outside that scope, and proving its absence requires inspecting the driver rather than trying more configuration knobs.
The kernel floor#
Kernel 7.1 is a hard floor if you want namespaces. On 7.0, rdma link add issued from inside a netns silently creates the device in init_net, and rdma dev set ... netns returns Operation not supported. The fix landed in commits 13f2a53c2a71 and f1327abd6abe, merged 2026-03-30 — after 7.0 closed. Ubuntu 26.04 ships 7.0, so this needs a mainline build.
rdma system set netns exclusive is not required: the upstream selftest tools/testing/selftests/rdma/rxe_rping_between_netns.sh never touches it, and everything here ran in shared mode.
What the negative result rests on#
All of this depends on not finding something, which makes the instrument the weak point rather than the result. Mine had three holes in it, in the order I found them:
A leftover process answered the test. An earlier run had failed and left its server
waiting in the background. The next run’s server printed rdma_bind_addr failed, and the
client connected to the stale one instead. The numbers were right by accident. Every run
after that started with an explicit pkill.
The capture was dropping two thirds of its packets. A broad udp port 4791 filter at
full rate lost 65 to 78 percent of what it saw to kernel drops, depending on the run. That
is tolerable when sampling TOS values across forty thousand packets. It is not tolerable
when the thing you are looking for is rare: CNP generation is rate-limited even on real
hardware, so a handful of missed packets is the difference between “no CNP” and a false
negative with four paragraphs of supporting evidence underneath it. Narrowing the filter so
the kernel discards non-matches before userspace, plus -s 128 -B 65536, took drops to zero.
The filter itself was never proven capable of catching a CNP. This is the one that mattered. I had validated that udp[8] == 0x11 matched thousands of real ACKs and concluded the byte offset was sound. But a filter that matches ACKs correctly can still be looking for the wrong opcode — and with the spec and Wireshark disagreeing on the CNP opcode, a wrong guess would have returned exactly the same zero. The fix was to inject synthetic CNPs at both values and confirm the capture caught them:
sudo pkill -f tcpdump; sleep 1
sudo tcpdump -i veth-b -nn -s 128 -B 65536 -w /tmp/cnp-probe.pcap \
'udp port 4791 and (udp[8] == 0x80 or udp[8] == 0x81)' &
sleep 1
sudo ip netns exec test1 python3 -c "
import socket
s = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
for op in (0x80, 0x81):
s.sendto(bytes([op]) + b'\x00'*31, ('172.16.0.2', 4791))"
sleep 1; sudo pkill -INT tcpdump; sleep 1
tcpdump -r /tmp/cnp-probe.pcap -nn | wc -l # must be 2That is a minimal packet, not a well-formed CNP. It does not need to be: the filter
inspects exactly one byte, udp[8], so a minimal packet is precisely what tests it.
Two injected, two captured. Only then does zero mean zero.
I was careful at L2, where being wrong would have cost an afternoon, and careless at L6, where being wrong would have meant publishing a negative result built on a filter nobody had tested. Closing it took two commands. I should have run them at the start.