From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: mail.toke.dk; dkim=pass header.d=google.com header.i=@google.com header.a=rsa-sha256 header.s=20251104 header.b=cqiBF5pc; arc=pass; dmarc=pass (Used From Domain Record) header.from=google.com policy.dmarc=reject Received: from mail-qk2-x11.google.com (mail-qk2-x11.google.com [IPv6:2607:f8b0:4864:34::11]) by mail.toke.dk (Postfix) with ESMTPS id 2EE7416F9B49 for ; Fri, 25 Sep 2026 15:03:31 +0200 (CEST) Received: by mail-qk2-x11.google.com with SMTP id d75a77b69052e-530d8a00bcdso6850861cf.2 for ; Fri, 25 Sep 2026 06:03:31 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1790341410; cv=none; d=google.com; s=arc-20260327; b=DY7zoHEHJVKmc90Gidx7jRUnDUwtsgaymvvaDr3PvzfqgwNGXrmJwOunoakTyhK7yW kldD8ejGPI179AA+WVdo56kX6idjN0RDNDE2uWq/YP+qb8R6oBAnoKF80l1MAdVKt7S+ qV7s0EtUM4KwrYy58S9UzTTpNeQPrGN1F52bCjrkcyeFCd2cSI8WHmFAV7GiNlBzBIpW Rh/xgxClGi4B9wu8rPCwL9ow+6NKws74ha5bslqwCK7LRnzUX7DlWv0vKl2HFVkLN8QU 5x0Dzhx1i0W3FtX1m8W6mCEVwYNwPyDaUV9R8xlK+N/6GJMMwFCCGgbEH/6i+aVjb6gp RTqg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:dkim-signature; bh=LO1mJQEWTKJy4KW79SYHwagOhWlUVH2rxVMnkykeDXc=; fh=hBXWMq6P6/35MEtwcQncHVoueM94QY6VGbnahi/WxH8=; b=bOAPxWLy+lB7pILFQi+Gq3FbaMh2QJ8pQkOLHyt7xQ/XQTiSv/eIRvm1RSGNIYSio7 bYShY9o6QjoIZhYIUkAXMFnB28cTVDvundlSVTUTaos1lcGHGCcNdy0IBMFrRmjzfAXf UWPlvyRV0s9ZidPnwFpvhyJPHIdc0vim2gz8Pa037pJIVCKAU9FjDJ91FfWCnNM6qrFL jV7S1kYTKBA0//YViOlDK2CQ+IWglioyi1yJJowP9m5VbVUFKaH0BcZNWqENgHAO67J6 bBM1bdvKHa9g1M5vJWChlhXNORdTbQYQGaNMfzwTJ6XsKheDb1quKYKFPrQeV/IfgsaI HSPw==; darn=lists.bufferbloat.net ARC-Authentication-Results: i=1; mx.google.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790341410; x=1790946210; darn=lists.bufferbloat.net; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:from:to:cc:subject :date:message-id:reply-to:content-type; bh=LO1mJQEWTKJy4KW79SYHwagOhWlUVH2rxVMnkykeDXc=; b=cqiBF5pcnqlOTYm4hm6UlrtwSf3+QPhjJ67hs7ChG2Kubaf0yzEDpPwrCSPOE/hNlZ pxHSfCLCNFn4gLvfoDP/obD81WDNYsobdLxC8P1+JkY3AeUd6D9WDuL7cF0qJx463fiO raxwdvXxP9KmHVFLorXkd1vHex7rlEPxQ0DNms+a/QZBL+3le3uhNPWfc0BWs84VhRAT PTLQCfYlFXr5y2Mz6t5PbapylwRz+tmlB52yUC9gdKiYcblyHl03lV3kP7Yw2wYiPqn0 bicwixOy4PM3qWJGC9R9Gn4yOx3QXjHxCVSuymjrNOQpwf2e3V/HY80xoukwrKVMzNAK Tp1Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790341410; x=1790946210; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=LO1mJQEWTKJy4KW79SYHwagOhWlUVH2rxVMnkykeDXc=; b=fTNigJehnPjRa50mnk2YI4PDqP9Us2jyO+vdw/Wdkk8q2W62Mvwm+jr81QTO4bxU21 SvSAHwAQb9tMiJCupf8YKTgNGQDOqIxCE4wE+2Ej78Od+xXKQCBe+UzQ0/jV4ffrRmTx OWW3wrFCxrQ+E2yJGO18TZcbgz2PFg+cKrMwf4IPNcODFzJjSe/TX2bqGIIiDhI9luah 6plLf+wgPyKsU9DZ9A7mMz/lW23NIhtRC4ASDGUVTmKVUPiGuKQ/TBRbrIUtDMeVyBJc RMJ9oZWbMoKNnOXNE1fZPbeARkgFYRggZq3MsFlrA3dtab/303AjK1joOsoB0yf+1zWR HzHA== X-Forwarded-Encrypted: i=1; AKwUvBzHROcIei9RSiI/Y3KCTEjtVLPbZpSHFiksGDC8qWFFelOo+kzzaNYe9Tu6wTv+VIGiPrlG@lists.bufferbloat.net X-Gm-Message-State: AFuF++nIabQohfcg15b/aUoQvJ/VSaMtuqZxNpxuUdwTZbBXVEwjevfg HcpJuljhdB571M99cVKtfwe0OBfc7mIPRQ6pYszcc5uVVTjRqbWsruqXkHZNg8H8Tf1Dra5cyzL 37z31ZUd+vzLn/aTZ7GIBYF/hAsk7cAKCGnPfB6vF X-Gm-Gg: AYBFou3LC1a9Ogwjd4MsabDlBHAgltczV1nIFjNCk3XCLJIgddo+KzBpzzEC1uZtS/V LwoxU9Ic3mQZIWXd6CRMFtNDUyb3VnM3gyD23Mriqlr1DzWoSwalP2zEci7jfljsCoQKU0N0qjc BnUyBr5ZkufdmR+cvv66LCyqDLJoY2EuDWNdTlYLcajZoa4beu/16xnp2TgkMWi/1QipBVSIsax o4dA05C12zsR6EGAiZHn6Oba2fv1jMvWJVrC1VfOvlLkMijZEzjsyG+TpfHZgY65rTWeF9DUlfF ERRG3iwH915rgvqGExLl+x60VFt/x1u9Lw5rhAcKkDSrvHuyVRsy/l6kmew+UUE/xWRY0ynvPmn hXgFbkfW19srVYiqsI1wueC6YCrQe2PJ8+1t0d21ISkRXP/dQ/Ihth0U5p3CTXj3oJktD X-Received: by 2002:ac8:7d12:0:b0:532:9c48:ec59 with SMTP id d75a77b69052e-5330b5d38dbmr44896871cf.18.1790341409488; Fri, 25 Sep 2026 06:03:29 -0700 (PDT) MIME-Version: 1.0 References: In-Reply-To: From: Eric Dumazet Date: Fri, 25 Sep 2026 15:03:18 +0200 X-Gm-Features: AclHuK-NpnkDVkhOVSxyv-HafVdDr-7CKXQ9rlYmLbL4LxpC386K5WFRODf_xz0 Message-ID: To: Jamal Hadi Salim Cc: netdev@vger.kernel.org, =?UTF-8?B?VG9rZSBIw7hpbGFuZC1Kw7hyZ2Vuc2Vu?= , cake@lists.bufferbloat.net, Jiri Pirko , "David S . Miller" , Jakub Kicinski , Paolo Abeni , Simon Horman , Victor Nogueira , hybris , Sashiko Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Message-ID-Hash: DKGGXNDLY4QR4TLDF2IVXSZYLY22HUI2 X-Message-ID-Hash: DKGGXNDLY4QR4TLDF2IVXSZYLY22HUI2 X-MailFrom: edumazet@google.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list Subject: [Cake] Re: [PATCH net-next] net/sched: fq_codel, cake: widen backlogs to u64 List-Id: Cake - FQ_codel the next generation Archived-At: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Fri, Sep 25, 2026 at 1:28=E2=80=AFPM Jamal Hadi Salim = wrote: > > On Fri, Sep 25, 2026 at 5:13=E2=80=AFAM Eric Dumazet wrote: > > > > On Fri, Sep 25, 2026 at 10:54=E2=80=AFAM Jamal Hadi Salim wrote: > > > > > > This is a follow-up to commit 8f735d64382d ("net/sched: bound > > > qdisc_pkt_len to prevent qdisc soft lockup"), which capped > > > qdisc_pkt_len() at QDISC_PKT_LEN_MAX (1 MiB). That cap bounds the sta= b > > > amplifier but leaves the per-flow backlog counter u32: > > > fq_codel_enqueue() accumulates qdisc_pkt_len(skb) into q->backlogs[id= x], > > > so a flow can still accumulate 4096 packets of 1 MiB each and wrap th= e > > > counter mod 2^32. After a wrap, fq_codel_drop() sees a tiny maxbacklo= g > > > and drops from an almost-empty flow, and the dequeue-side subtraction= s > > > corrupt the counter further. > > > > > > Widen the fq_codel backlogs table, the fat-flow scan (maxbacklog/len)= and > > > the drop threshold to u64. fq_codel is not lockless: every writer run= s > > > under the root qdisc lock, so plain u64 arithmetic keeps the WRITE_ON= CE > > > publish / READ_ONCE-consume pattern. The dump path > > > (fq_codel_dump_class_stats) stays a lockless stat-only read. > > > > > > CAKE accumulates the same generic qdisc_pkt_len(skb) into its per-flo= w > > > b->backlogs[] and per-tin b->tin_backlog and consumes the values for > > > longest-flow pruning (cake_heapify/cake_heapify_up) and for the shape= r > > > staleness check, so it shares the bug. Widen those counters and the h= eap > > > comparison locals to u64; the class/tin stats keep exporting the low = 32 > > > bits through the unchanged uAPI fields. > > > > > > Conditions to recreate the bug: CAP_NET_ADMIN in a user namespace; > > > CONFIG_NET_SCH_FQ_CODEL=3Dy. > > > > > > ip tuntap add tun0 mode tun > > > ip link set tun0 txqueuelen 32 up > > > ip addr add 10.99.0.1/24 dev tun0 > > > tc qdisc add dev tun0 root handle 1: stab overhead 2000000000 \ > > > fq_codel flows 1 limit 4200 ecn drop_batch 4096 > > > # hold the tun fd open without reading (IFF_BACKPRESSURE) so the qd= isc > > > # backlog persists, then send at least 4300 packets (the wrap start= s > > > # at 4096 resident; the over-limit drop that reads the wrapped > > > # threshold fires past the 4200 limit): backlogs[0] wraps at 4096 x > > > # 1 MiB and the fat-flow threshold reads the wrapped value. > > > > > > With a 1 MiB qdisc_pkt_len cap the counter wraps at 4096 resident > > > packets. At limit 4200 the first over-limit enqueue (the 4201st) sees= a > > > wrapped 105 MiB (half-backlog threshold 52 MiB, a ~52 packet drop bur= st), > > > where the u64 counter sees 4201 MiB (threshold 2100 MiB, a ~2100 pack= et > > > drop burst). > > > > > > Reported-by: Sashiko (nipa) > > > Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/202608181= 01130.16203-1-jhs@mojatatu.com > > > Link: https://lore.kernel.org/netdev/20260818101130.16203-1-jhs@mojat= atu.com/ > > > Tested-by: hybris > > > Signed-off-by: Jamal Hadi Salim > > > --- > > > > This makes no sense. > > > > As absurd as it looks that code is reachable ;-> > > > These qdisc have been developped to address bufferbloat issues. > > > > Storing 4GB in a qdisc is absolutely insane. > > > > The counter wrap is not because we stored 4GB, rather it is because > "tc .. stab overhead ..." inflates qdisc_pkt_len() for a 64B pkt to > 1MB. So ~4K packets (put in other words a few "real" KB) makes that > backlog[0] cross 2^32. Kill this stuff ? Who is still using this, for what reason ? Really, it is time we stop adding code only for fuzzers. > Result is pruning the wrong flow.. > > > Let's drop at enqueue if the current backlog is approaching 4GB (this > > can later be a new config/attribute in net-next) > > Note, this is net-next already. Enqueue is fast path - are you ok with th= at? > > cheers, > jamal