darrenqu.net

AI Networking

RoCE Congestion Control Without RDMA NICs: Soft-RoCE Gives You One of DCQCN's Three Roles

4119 words 20 min read

roce

tc marks 3,682 packets CE. All 3,682 arrive at the receiver. Soft-RoCE sends back zero CNPs. Where the line falls between what you can verify in software and what needs hardware — and how I nearly got the answer wrong.

on this page

Every explanation of DCQCN assumes you have RDMA NICs. Every soft-RoCE guide stops at “ib_write_bw returns a number.” I could not find anyone who joins the two up, so the question goes unanswered: without RDMA hardware, how much of RoCE congestion control can you actually verify?

DCQCN has three roles#

The feedback loop needs three participants:

RoleWho plays itWhat it does
CP — Congestion Pointthe switchmarks the IP header ECN field as CE when the queue builds
NP — Notification Pointthe receiving NICsees CE, sends a CNP back to the sender
RP — Reaction Pointthe sending NICreceives the CNP, cuts the rate for that QP

The CP runs on a switch, but the NP and RP run inside NIC hardware. DCQCN’s rate control is a per-QP state machine with its own timers and recovery curve, implemented in silicon — not something a driver computes in software.

The lab, and which parts of it are honest#

One VM. One veth pair, one end moved into a network namespace. rdma_rxe on both ends, tc in the middle.

netns test1                          init_net (host)
┌────────────────┐                  ┌────────────────┐
│ veth-a   rxe0  │══════════════════│ veth-b   rxe1  │
│ 172.16.0.1/24  │   tbf + red ecn  │ 172.16.0.2/24  │
└────────────────┘                  └────────────────┘
      NP role             CP role         RP role

Before running anything, define what each piece represents:

ComponentStands in forFaithful?
rdma_rxe (soft-RoCE)an RDMA NICdata transport yes, hardware no. The RoCEv2 data and ACK wire formats are exercised here, but there is no kernel bypass, no CNP generation and no DCQCN state machine
veth paira cableno real buffer, no DCB, no PFC
tc red ... ecna switch’s WRED + ECN markingfaithful for the marking behavior tested here. Linux RED genuinely rewrites the ECN field on queue build-up, but it is not a physical switch buffering model

The ECN-marking function playing the congestion point is real. The parts playing the NICs are not. That determines what the lab can prove and the order of the tests below.

Lab scope. Ubuntu 26.04, kernel 7.2.6-070206-generic, rdma-core 61.0-2ubuntu3, single VM, 4 vCPU / 2 GB. Soft-RoCE only — no RDMA hardware was involved, and nothing here should be extrapolated to real NIC behavior.

Building it#

Everything below runs on a single VM. No RDMA hardware, no second machine. The one read-only checks in the first block need no privileges; everything after them does.

Check the kernel version first. It is the cheapest gate and it decides whether any of the rest is possible — on 7.0 the topology below builds without error and then quietly does the wrong thing, putting both RDMA devices in init_net:

uname -r                                   # needs 7.1 or newer
grep RDMA_RXE /boot/config-$(uname -r)     # expect CONFIG_RDMA_RXE=m

Ubuntu 26.04 ships 7.0, which is not enough; everything here ran on a mainline build, 7.2.6-070206-generic. The reason 7.1 is the floor is at the end of the article.

sudo modprobe rdma_rxe

sudo ip netns add test1
sudo ip link add veth-a type veth peer name veth-b
sudo ip link set veth-a netns test1

sudo ip netns exec test1 ip addr add 172.16.0.1/24 dev veth-a
sudo ip netns exec test1 ip link set veth-a up
sudo ip netns exec test1 ip link set lo up
sudo ip addr add 172.16.0.2/24 dev veth-b
sudo ip link set veth-b up

sudo ip netns exec test1 rdma link add rxe0 type rxe netdev veth-a
sudo rdma link add rxe1 type rxe netdev veth-b

Now confirm the namespace really got its own RDMA stack. Both sides must show a listener on 4791: that per-namespace UDP socket is precisely what Linux 7.1 added, so this is the runtime proof, where uname -r was only the claim. A vendor kernel can report a new version without carrying the patch.

sudo ip netns exec test1 ss -Huln 'sport == :4791'   # must not be empty
sudo ss -Huln 'sport == :4791'
sudo ip netns exec test1 rdma link                   # must show 'netdev veth-a'
sudo ip netns exec test1 ping -c 2 -W 1 172.16.0.2

If the namespace side is empty, or rdma link there shows nothing while the host shows both devices, stop — the kernel is too old and the rest will not mean anything.

Then generate traffic. -R uses rdma_cm for connection setup, which selects the RoCEv2 GID automatically; --tos=2 sets DSCP 0 with ECT(0), which is what makes RED mark rather than drop:

sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1

Start every run by killing the previous one. A server left waiting from a failed attempt will happily answer the next client that comes along, and you will be measuring something other than what you think — more on that at the end.

sudo pkill -f ib_write_bw; sudo pkill -f tcpdump

That is all you need between runs. The pkill -f above matches every ib_write_bw and tcpdump process on the VM, so use it only on a dedicated lab VM; on a shared host, record the process IDs and stop only what this experiment started.

Everything from here to the end of the article assumes the topology is still up, so leave it in place until you are finished. The block below removes it:

sudo pkill -f ib_write_bw; sudo pkill -f tcpdump
sudo tc qdisc del dev veth-b root
sudo ip netns exec test1 rdma link del rxe0
sudo rdma link del rxe1
sudo ip link del veth-b
sudo ip netns del test1

A reboot has the same effect: namespaces, veth pairs and rxe devices are all gone afterwards. Re-run the build block above and carry on.

Verify the instrument before the subject#

The rule I followed, learned the hard way on an earlier EVPN lab: check that the measurement works before trusting what it says. Each layer is a precondition for the next.

L1 RDMA works → L2 it really is RoCEv2 on the wire → L3 DSCP is settable → L4 packets are ECN-capable → L5 CE marks get produced → L6 CNPs come back → L7 the rate drops.

Layer 3 is easy to skip. Setting a DSCP value and reading it back is not a test of RoCE; it verifies the capture path before later results depend on it.

L1–L5: everything up to and including the switch works#

L1. ib_write_bw between the two devices: 527.37 MiB/sec. Soft-RoCE is CPU-bound — at full speed exactly three threads saturate, the client, the server, and the rxe kworker. More cores only move this number; the L5 and L6 runs are shaped to 100 Mbit and barely register, so a smaller VM reaches the same conclusions with a different L1 figure.

 #bytes     #iterations    BW peak[MiB/sec]    BW average[MiB/sec]   MsgRate[Mpps]
 65536      5000             527.37             485.77               0.007772

L2. Capture the exchange and look at what is actually on the wire:

mkdir -p ~/roce-lab/evidence/
sudo tcpdump -i veth-b -nn -U -w ~/roce-lab/evidence/L2-roce-v2.pcap udp port 4791 &
sleep 1
sudo ip netns exec test1 ib_write_bw -d rxe0 -R >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R -n 5000 172.16.0.1 | tail -3
sudo pkill -INT tcpdump

261,403 packets, every one of them on UDP 4791. Three dominant length clusters account for 261,385 of them:

CountUDP lengthBreakdownType
253,314104012 (BTH) + 1024 + 4 (ICRC)RDMA Write Middle/Last
4,025105612 (BTH) + 16 (RETH) + 1024 + 4RDMA Write First
4,0462012 (BTH) + 4 (AETH) + 4 (ICRC)Mostly Acknowledge

The header sizes are fixed by the specification, and only one assignment makes the bytes add up — RETH carries the remote address, rkey and length, so it can only ride on a First. That is arithmetic, though, not measurement. Counting opcodes gives a second, independent answer:

for op in 0x06 0x07 0x08 0x11; do
  printf 'opcode %-5s : %8s\n' "$op" \
    "$(tcpdump -r ~/roce-lab/evidence/L2-roce-v2.pcap -nn "udp[8] == $op" 2>/dev/null | wc -l)"
done
opcode 0x06  :     4025      RDMA Write First
opcode 0x07  :   249292      RDMA Write Middle
opcode 0x08  :     4022      RDMA Write Last
opcode 0x11  :     4040      Acknowledge

The data-packet counts line up to the packet. 4,025 Firsts is exactly the 1056-byte cluster. 249,292 Middles plus 4,022 Lasts is 253,314 — exactly the 1040-byte cluster. The 20-byte cluster is mostly Acknowledges: 4,040 of 4,046 carry opcode 0x11. Six carry a different opcode, and another 18 packets fall outside the three length clusters above. I did not chase those 24 remaining packets further.

65536 ÷ 1024 bytes of RDMA payload = 64 packets per message, and the opcodes show how those 64 divide: one First carrying the RETH, then 62 Middles, then one Last. 249,292 ÷ 4,025 = 61.94, and First, Last and Acknowledge agree with one another to within half a percent — one of each per message. The residual differences are small enough to be capture boundaries rather than protocol behavior.

L3 and L4. perftest’s -T/--tos sets the TOS byte via rdma_cm. The open question was whether the low two bits survive — many stacks mask the ECN field out of a user-supplied TOS, on the grounds that ECN belongs to the transport. Run it four times, changing only --tos, and read the byte back off the wire each time:

for T in "" 104 2 106; do
  sudo pkill -f ib_write_bw >/dev/null 2>&1; sudo pkill -f tcpdump >/dev/null 2>&1; sleep 1
  F=~/roce-lab/evidence/tos-${T:-unset}.pcap
  sudo tcpdump -i veth-b -nn -U -w "$F" udp port 4791 & sleep 1
  sudo ip netns exec test1 ib_write_bw -d rxe0 -R ${T:+--tos=$T} >/dev/null 2>&1 & sleep 2
  sudo ib_write_bw -d rxe1 -R ${T:+--tos=$T} -n 2000 172.16.0.1 >/dev/null 2>&1
  sudo pkill -INT tcpdump; sleep 1
  printf '=== --tos=%s ===\n' "${T:-unset}"
  tcpdump -r "$F" -nn -v 2>/dev/null | grep -oE 'tos 0x[0-9a-f]+' | sort | uniq -c
done

They do not get masked:

--tosobserved on the wirepacketsDSCPECN
unset0x028,599000 Not-ECT
1040x6832,1902600
20x231,501010 ECT(0)
1060x6a30,8162610 ECT(0)

The packet counts differ between runs only because a broad capture at full rate drops most of what it sees — 75 to 78 percent in these four. The TOS value is the measurement; the volume is not.

RED drops Not-ECT packets rather than marking them, so without ECT there is no L5. Because rdma_cm passes the ECN bits through, no packet forging is needed.

One structural detail: in each of the three runs where the data carries a non-zero TOS, exactly 39 packets still read tos 0x0 — the same 39 every time. They are consistent with the rdma_cm exchange, which happens before the QP exists and therefore before the TOS option can apply to it. In the unset run there is no separate 39, because the exchange and the data are both 0x0 and cannot be told apart. Confirming the opcodes and sequence of those 39 packets would turn that inference into direct evidence.

L5. Shape well below the ~485 MiB/s the unshaped run managed — call it 4 Gbit/s — so that a queue actually forms:

sudo tc qdisc add dev veth-b root handle 1: tbf rate 100mbit burst 32kb latency 50ms
sudo tc qdisc add dev veth-b parent 1:1 handle 10: red \
   limit 400000 min 30000 max 100000 avpkt 1000 burst 55 ecn bandwidth 100mbit

Omitting bandwidth is a quiet trap — RED prints RED: set bandwidth to 10Mbit and computes its averaging constants against the wrong link rate.

Now run traffic through it and read the qdisc counters. Unlike a packet capture, they are not affected by tcpdump drops:

sudo pkill -f ib_write_bw; sudo pkill -f tcpdump; sleep 1

sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1 | tail -3

sudo tc -s qdisc show dev veth-b

Shaped to 100 Mbit, 5000 iterations of 64 KB take about half a minute instead of the second they took unshaped. The queue must build before RED has anything to act on.

qdisc tbf 1: root refcnt 129 rate 100Mbit burst 32Kb lat 50ms
 Sent 346324396 bytes 320040 pkt (dropped 0, overlimits 834591 requeues 0)
qdisc red 10: parent 1:1 limit 400000b min 30000b max 100000b ecn
 Sent 346324396 bytes 320040 pkt (dropped 0, overlimits 3698 requeues 0)
  marked 3698 early 0 pdrop 0 other 0

Read all three numbers together. marked means CE marks were produced. early 0 pdrop 0 means nothing was dropped — proof that ECT(0) took effect, because RED recognized the packets as ECN-capable and marked them instead.

marked is a counter on the sending qdisc. It proves marks were applied, not that they survived the veth traversal and reached the receiver. So measure inside the namespace, on the ingress side:

sudo pkill -f ib_write_bw; sudo pkill -f tcpdump; sleep 1
sudo tc -s qdisc show dev veth-b | grep marked          # baseline, counters are cumulative

sudo ip netns exec test1 tcpdump -i veth-a -nn -s 128 -B 65536 \
  -w ~/roce-lab/evidence/ce-arrival.pcap 'udp port 4791 and (ip[1] & 0x03) == 3' &
sleep 1

sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1 | tail -3
sudo pkill -INT tcpdump; sleep 1

sudo tc -s qdisc show dev veth-b | grep marked          # subtract the baseline
tcpdump -r ~/roce-lab/evidence/ce-arrival.pcap -nn | wc -l

The capture runs inside the namespace on veth-a, the receiving side of the link. The filter (ip[1] & 0x03) == 3 selects packets whose IP header carries ECN 11 — Congestion Experienced.

Measurement pointValue
qdisc marked, before this run3,698
qdisc marked, after7,380
delta — marks applied by the sender3,682
CE-marked packets received inside the namespace3,682

An exact match, and the capture that produced the second figure dropped nothing: 3682 packets captured / 3682 received by filter / 0 dropped by kernel. Every mark RED applied on the way out arrived at the far end still carrying it. The congestion signal demonstrably reached rxe0.

The CP’s ECN-marking function is reproducible in software. Linux tc acts on the same IP ECN field, while its queue and buffer model remains different from a physical switch ASIC.

L6: 3,682 CE marks in, zero CNPs out#

Capturing the reverse direction needs a narrow filter and a small snaplen, or the kernel starts dropping — which for a rare packet would be fatal. udp[8] is the first byte of the UDP payload, which is the BTH opcode:

sudo pkill -f ib_write_bw; sudo pkill -f tcpdump; sleep 1
rm -f ~/roce-lab/evidence/L6-reverse.pcap        # never census a capture from an older run

sudo tcpdump -i veth-b -nn -s 128 -B 65536 \
  -w ~/roce-lab/evidence/L6-reverse.pcap 'udp port 4791 and src 172.16.0.1' &
sleep 1

sudo ip netns exec test1 ib_write_bw -d rxe0 -R --tos=2 >/dev/null 2>&1 &
sleep 2
sudo ib_write_bw -d rxe1 -R --tos=2 -n 5000 172.16.0.1 | tail -3
sudo pkill -INT tcpdump; sleep 1

tcpdump -r ~/roce-lab/evidence/L6-reverse.pcap -nn | wc -l
for i in $(seq 0 255); do
  op=$(printf '0x%02x' $i)
  c=$(tcpdump -r ~/roce-lab/evidence/L6-reverse.pcap -nn "udp[8] == $op" 2>/dev/null | wc -l)
  [ "$c" -gt 0 ] && printf 'opcode %-5s : %s\n' "$op" "$c"
done

And back the other way. 5,036 packets matched the filter and 5,020 reached the file — the sixteen in between were still in tcpdump’s buffer when it was signaled, and the kernel dropped none at all. All 256 opcode values were then swept across those 5,020:

BTH opcodeMeaningCount
0x11Acknowledge5,009
0x04RC SEND Only (perftest sync)9
0x64UD SEND Only (rdma_cm control)2
0x80 / 0x81CNP0

5,020 of 5,020 accounted for, with nothing left over.

CNP is opcode 0x81. Wireshark’s InfiniBand dissector disagrees: its bth_opcode_tbl maps 0x80 to CNP and has no entry at 0x81, with the opcode field declared unmasked. A real CNP therefore appears as an unnamed opcode 129, and filtering on the display name finds nothing. The specification (0b10000001), the IETF drafts, NVIDIA’s documentation (decimal 129, “RoCEv2 Congestion Management Ack”), DPDK’s BTH flow-matching code and Scapy’s RoCE module all say 0x81. Scapy provides a quick check — python3 -c "from scapy.contrib.roce import CNP_OPCODE; print(hex(CNP_OPCODE))":

CNP_OPCODE = 0x81

def cnp(dqpn):
    return BTH(opcode=CNP_OPCODE, becn=1, dqpn=dqpn) / CNPPadding()

DPDK’s patch is explicit about exactly this use case:

flow create 0 group 1 ingress pattern eth / ipv4 / udp dst is 4791 /
  ib_bth opcode is 0x81 dst_qp is 0xd3 / end actions queue index 0 / end

I swept both values regardless. Zero at each.

Wireshark has a second blind spot in the same area, and this one bites anyone doing congestion work on real hardware too. Here is a real RDMA WRITE Middle packet from this lab, dissected in full:

InfiniBand
    Base Transport Header
        Opcode: Reliable Connection (RC) - RDMA WRITE Middle (7)
        0... .... = Solicited Event: False
        .0.. .... = MigReq: False
        ..00 .... = Pad Count: 0
        .... 0000 = Header Version: 0
        Partition Key: 65535
        Reserved: 00
        Destination Queue Pair: 0x000022
        0... .... = Acknowledge Request: False
        .000 0000 = Reserved (7 bits): 0
        Packet Sequence Number: 16516294

Notice how much of that is broken out to individual bits. Solicited Event, MigReq, Pad Count and Header Version each get their own bitmask line; so does Acknowledge Request, and even the seven reserved bits beside it. Then look at the line that just says Reserved: 00.

That byte is not reserved. Bits 7 and 6 are FECN and BECN, InfiniBand’s own congestion notification bits — the second of the two congestion mechanisms RoCE carries. The kernel knows they are there; rxe_hdr.h defines BTH_FECN_MASK (0x80000000) and BTH_BECN_MASK (0x40000000) against exactly this word. Wireshark does not decode them at all: the strings fecn and becn appear nowhere in its InfiniBand dissector, and the byte is added to the tree as an undifferentiated Reserved. There is no field to filter on.

To see those two bits you have to read the byte yourself — udp[12] & 0x80 for FECN, udp[12] & 0x40 for BECN. In this capture the byte reads 00, which is what the driver source predicts: those accessors exist and nothing ever calls them.

The receiver’s counters provide another check. rdma statistic show link rxe0/1 exposes eighteen counters — packets, bytes, sequence errors, RNR errors, retries — and not one of them relates to ECN, CE or CNP. rxe does not merely fail to respond to congestion marks; it does not count them.

And the driver source says why. rxe_resp.c, the responder path where an NP role would live, never reads the IP header ECN bits and never generates a CNP. The InfiniBand BTH carries its own congestion bits, FECN and BECN, and rxe does define accessors for them in rxe_hdr.h — bth_fecn(), bth_becn(), bth_set_fecn(), bth_set_becn(). Checking the 37 source files in drivers/infiniband/sw/rxe/, those accessors are referenced zero times outside their own definitions. They are inherited header plumbing that nothing calls.

The tested kernel and the mainline 7.3-rc source snapshot checked for this article both lack these paths.

Where the line falls#

Verifiable in software: the RoCEv2 data and ACK wire formats exercised here, DSCP marking, ECT(0) transmission, and CE-mark generation and delivery. L1 through L5, with a complete evidence chain.

Not verifiable: the other two roles of the DCQCN loop — the receiver returning a CNP, and the sender reacting to it.

Why: those roles are implemented in NIC hardware. In kernel 7.2.6 and the mainline 7.3-rc snapshot examined here, Soft-RoCE implements the data transport but not the DCQCN congestion-control paths.

The boundary is not simply hardware versus software; it is between IP-layer marking and RoCE transport behavior.

ECN lives in the IP header — two bits, defined by RFC 3168. Marking CE is an IP-layer action, which is why Linux tc can reproduce the marking behavior tested here while acting on a field in a layer it already owns.

CNP lives one layer up, in the RoCE transport. Generating one means reading a field in the IP header and then emitting a new RoCE transport packet carrying a particular BTH opcode. That crossing — IP signal in, RoCE transport packet out — is what the examined rxe versions do not implement. The IP-header marking path tested here reproduces in software. In the RoCE transport, soft-RoCE gives you the data format but not the congestion-control behavior.

PFC is further out of reach still. veth has no DCB and no priority queues, so 802.1Qbb cannot be approached at all.

A richer simulator would have to model the NIC behavior. NVIDIA’s DSX Air markets “emulated hardware fidelity to the network adapter” and “simulation of NVIDIA ConnectX SuperNICs”. On a standard Ubuntu host node there, lspci returns QEMU devices and nothing else.

00:00.0 Host bridge: Intel Corporation 440FX - 82441FX PMC [Natoma]
00:08.0 Ethernet controller: Red Hat, Inc. Virtio network device [1af4:1000]
...
driver: virtio_net

No Mellanox device, no mlx5. RoCE on that node uses soft-RoCE again, with the same missing paths. I did not survey every node type DSX Air offers, and the documentation has no page describing what is and is not simulated, so treat this as one data point rather than a verdict on the product.

Without hardware, you can still practise RoCE protocol structure, QoS marking and switch-side ECN configuration on one small VM. The DCQCN feedback loop is outside that scope, and proving its absence requires inspecting the driver rather than trying more configuration knobs.

The kernel floor#

Kernel 7.1 is a hard floor if you want namespaces. On 7.0, rdma link add issued from inside a netns silently creates the device in init_net, and rdma dev set ... netns returns Operation not supported. The fix landed in commits 13f2a53c2a71 and f1327abd6abe, merged 2026-03-30 — after 7.0 closed. Ubuntu 26.04 ships 7.0, so this needs a mainline build.

rdma system set netns exclusive is not required: the upstream selftest tools/testing/selftests/rdma/rxe_rping_between_netns.sh never touches it, and everything here ran in shared mode.

What the negative result rests on#

All of this depends on not finding something, which makes the instrument the weak point rather than the result. Mine had three holes in it, in the order I found them:

A leftover process answered the test. An earlier run had failed and left its server waiting in the background. The next run’s server printed rdma_bind_addr failed, and the client connected to the stale one instead. The numbers were right by accident. Every run after that started with an explicit pkill.

The capture was dropping two thirds of its packets. A broad udp port 4791 filter at full rate lost 65 to 78 percent of what it saw to kernel drops, depending on the run. That is tolerable when sampling TOS values across forty thousand packets. It is not tolerable when the thing you are looking for is rare: CNP generation is rate-limited even on real hardware, so a handful of missed packets is the difference between “no CNP” and a false negative with four paragraphs of supporting evidence underneath it. Narrowing the filter so the kernel discards non-matches before userspace, plus -s 128 -B 65536, took drops to zero.

The filter itself was never proven capable of catching a CNP. This is the one that mattered. I had validated that udp[8] == 0x11 matched thousands of real ACKs and concluded the byte offset was sound. But a filter that matches ACKs correctly can still be looking for the wrong opcode — and with the spec and Wireshark disagreeing on the CNP opcode, a wrong guess would have returned exactly the same zero. The fix was to inject synthetic CNPs at both values and confirm the capture caught them:

sudo pkill -f tcpdump; sleep 1
sudo tcpdump -i veth-b -nn -s 128 -B 65536 -w /tmp/cnp-probe.pcap \
  'udp port 4791 and (udp[8] == 0x80 or udp[8] == 0x81)' &
sleep 1

sudo ip netns exec test1 python3 -c "
import socket
s = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
for op in (0x80, 0x81):
    s.sendto(bytes([op]) + b'\x00'*31, ('172.16.0.2', 4791))"

sleep 1; sudo pkill -INT tcpdump; sleep 1
tcpdump -r /tmp/cnp-probe.pcap -nn | wc -l          # must be 2

That is a minimal packet, not a well-formed CNP. It does not need to be: the filter inspects exactly one byte, udp[8], so a minimal packet is precisely what tests it.

Two injected, two captured. Only then does zero mean zero.

I was careful at L2, where being wrong would have cost an afternoon, and careless at L6, where being wrong would have meant publishing a negative result built on a filter nobody had tested. Closing it took two commands. I should have run them at the start.

← more in AI Networking