From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: mail.toke.dk; dkim=pass header.d=google.com header.i=@google.com header.a=rsa-sha256 header.s=20251104 header.b=vuZUnMjx; arc=pass; dmarc=pass (Used From Domain Record) header.from=google.com policy.dmarc=reject Received: from mail-qk2-x0f.google.com (mail-qk2-x0f.google.com [IPv6:2607:f8b0:4864:34::f]) by mail.toke.dk (Postfix) with ESMTPS id 25F8516F9F21 for ; Fri, 25 Sep 2026 15:43:47 +0200 (CEST) Received: by mail-qk2-x0f.google.com with SMTP id d75a77b69052e-530de452c33so9076041cf.1 for ; Fri, 25 Sep 2026 06:43:46 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1790343825; cv=none; d=google.com; s=arc-20260327; b=T6SiF6GXpNoZRSA2uVVRCtMIk63jol7aVbqIbkjR6Kl8KIPxKnZlDIQGpJpv8TZ8MS Zvy8wHAXDBTion9FxTQmHzILoBHjXofJNQ2rId26yzlowHlP1/WeaOVe/toiPMzL+NIK pzTt38Jl1hRV0eTfrbW6ZIhdXb/VjVhraLXl1KMrErVyHMux/06fsc5ijyonLer3BpNJ 09nYXqUHU8XW1/U1AsCTD4UOZn+eWnj3fCRsgombRqoD/YOxzYbZHUmQEaSkgi1w9ekj KTdGcoV2dePJDcvarrPasRHN5V0fk+5rV3cCQdsmH/TPkYhRidUtx+AaKJer598tFT9H hwxQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:dkim-signature; bh=aS8QAswde3y4iU7Q8F3Yth5T0bAE/PRbpFF/5Gg231k=; fh=xQVCcuaKrpo260S6sqvdVLHD5EtWHltY8mVs4E9iX2w=; b=WB5FvjiIhYCrnURZrVNrM0WIo+AvDyQLXNjB7fs6iicGYvNUdOsyIpq4NIzOJ1Qw5B p4iV39Ia4Qmtlg4w5OJ3QJoVMCT3Uw5knGtVMT/6eRr6alR1OMCKGSDHeRD50orNNStv bWMDR9dh4HpjQYnVo3h9kxrtnlv/45udPie1RRsT4GGs+XDz2VACc0ARdL7PMf47Gm6m mSrHQu2lZ4t3erMSGLmlA+TW8uO4vc6Q713bPipjX3QNsrn8U1cGaDbSZE5RzKp8IrOt XSB7GOjhOkWYOihhaTrIOowEfp3e+jL4toiOkLUf+YGqGEWune8tZOX2Y9IixBLZbwUq E+1g==; darn=lists.bufferbloat.net ARC-Authentication-Results: i=1; mx.google.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790343825; x=1790948625; darn=lists.bufferbloat.net; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:from:to:cc:subject :date:message-id:reply-to:content-type; bh=aS8QAswde3y4iU7Q8F3Yth5T0bAE/PRbpFF/5Gg231k=; b=vuZUnMjx9ElS7CuKWpbXglTu3h3052rcrBpz+KJGGDBLaXnlqE6uT2mtf2gyMOoQB6 aIcu36n2JnRAaJGAZOKhfv3r739E2aTeUFNG0yhpFUNsXffvvLdJCSpoLV/9J48UA478 lDn7DaffzcfHEOpSZ5BlOyhTJosdt8m0Y5qJNjI9ZOdnDfrmvFJ7rP2DU2IpfpCg3V3P NWlDxTqhRFfpa46DdSa4qEcwURMHUiCMATLfwZ9ZG+tYd9BpMlGrWt9nVPxVCn9aWog7 r5tfHp3bSga3F7jUr0t8B1jSfufPHbqWXVw+9jn349Z7M9zyIgoz2jy8WepLkASHIz1A mf+g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790343825; x=1790948625; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=aS8QAswde3y4iU7Q8F3Yth5T0bAE/PRbpFF/5Gg231k=; b=YG2zd8ExmYknfqHIYiXeRbg+NEsrohldBV2R4nuOUAFgntkgNcybXAO8g/tmF0R35m RfHGyqtxjYij4e/kU+GdPxMh5TAT9P7sxQr4oY055X5TBWPrakFDWcyTEPAmsa8ELZIn 7Si8cVZ/UX+fjNlGm58Vk2DZv1Noh1XmfinxrY50TJWysIaam0gDSA3iLlosy/5wdnPC oY1nAOCBFZ6vsl3U2DgNfkYat2rKytYUupvdHLEiuDQL+vX9X8NLPWIy5GVg3EZI4R4j GU/HnRGwi2giA6IIPQWx9tK8vpta245BJ2A9qyse/9HVAa9XEdHV2fvJgvCL/3qQ/bCa JX2g== X-Forwarded-Encrypted: i=1; AKwUvBx9cPwtTKGlFyz0ke7GA9a8xXzLA5EwLMWi6hOSV+i0EPtgx6x+FqQlnXZsESg4cV6lEWZU@lists.bufferbloat.net X-Gm-Message-State: AFuF++lwnwsviQuVXHuWJG79QO1BRoJ9t8zNTtv+SsgII9hh4Fhn+4sS /90YtmRyZaf77Jlcc8bUBvrnuJh997ojWrLX4jPY203V6LGSFNF6dd7Jf5IxiM0eEe+ti+tuPbh vUBeuF2AAOIxSl49t5WCVJ2Do0ugcR40RSR0orP2F X-Gm-Gg: AYBFou3UOA83TMFQP72M3gKMNe8Gx7eI4TBZBNuidaHd8YIqG1pxNXdaB7QvU/WalHu oyMHeeWfkDjAQ/YoArWeIcXKRkvcwqn6qUMsq4acJwE5kW1L1/iHgOSYopU9TqNmW4o4idtMHe/ IqcQMqPlZNWRPx/m51Uv/N8O070Z87QdBgO8ISUbyUfHoLpNNmloD7lqyc9eDL4xxrC4O6k9+Gn GgYCmQ7XdoC2GO0vX+Hfajk0c9ouZ0aQ05XRpMa0E75ZHcbB2QdTuhVa6CcGfNn7mEeSlqUTpU/ yOHcOl4S6Bwkp0dY8yD75n/d5Ejs9NvqhaGWkZKGwFI1fr9ItcTDoAs73S5wLOerRA1HhSY/e2Z ee28wnQkBtguLfKW6fvuumCA/UuSrDnqHiB45NuycCSjEEMre3z/x8HTOlx1zyXs2wgp7 X-Received: by 2002:ac8:5f52:0:b0:530:b2e2:903c with SMTP id d75a77b69052e-5330b705d24mr42685811cf.60.1790343823975; Fri, 25 Sep 2026 06:43:43 -0700 (PDT) MIME-Version: 1.0 References: <68BE1514-828D-4184-84E4-90F2EB3F035D@gmx.de> In-Reply-To: <68BE1514-828D-4184-84E4-90F2EB3F035D@gmx.de> From: Eric Dumazet Date: Fri, 25 Sep 2026 15:43:32 +0200 X-Gm-Features: AclHuK9vscudfvim8sMExKqr7dvfM9Zf6NqeO_k_orfeFHTKqq4bkpBTUJtLB8U Message-ID: To: Sebastian Moeller Cc: Jamal Hadi Salim , netdev@vger.kernel.org, =?UTF-8?B?VG9rZSBIw7hpbGFuZC1Kw7hyZ2Vuc2Vu?= , cake@lists.bufferbloat.net, Jiri Pirko , "David S . Miller" , Jakub Kicinski , Paolo Abeni , Simon Horman , Victor Nogueira , hybris , Sashiko Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Message-ID-Hash: EL7H2GBITUVARKPRLMHZGAMAJ4PJCLLZ X-Message-ID-Hash: EL7H2GBITUVARKPRLMHZGAMAJ4PJCLLZ X-MailFrom: edumazet@google.com X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation; nonmember-moderation; administrivia; implicit-dest; max-recipients; max-size; news-moderation; no-subject; digests; suspicious-header X-Mailman-Version: 3.3.10 Precedence: list Subject: [Cake] Re: [PATCH net-next] net/sched: fq_codel, cake: widen backlogs to u64 List-Id: Cake - FQ_codel the next generation Archived-At: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: On Fri, Sep 25, 2026 at 3:30=E2=80=AFPM Sebastian Moeller = wrote: > > Hi Eric, > > > > On Sep 25, 2026, at 15:03, Eric Dumazet via Cake wrote: > > > > On Fri, Sep 25, 2026 at 1:28=E2=80=AFPM Jamal Hadi Salim wrote: > >> > >> On Fri, Sep 25, 2026 at 5:13=E2=80=AFAM Eric Dumazet wrote: > >>> > >>> On Fri, Sep 25, 2026 at 10:54=E2=80=AFAM Jamal Hadi Salim wrote: > >>>> > >>>> This is a follow-up to commit 8f735d64382d ("net/sched: bound > >>>> qdisc_pkt_len to prevent qdisc soft lockup"), which capped > >>>> qdisc_pkt_len() at QDISC_PKT_LEN_MAX (1 MiB). That cap bounds the st= ab > >>>> amplifier but leaves the per-flow backlog counter u32: > >>>> fq_codel_enqueue() accumulates qdisc_pkt_len(skb) into q->backlogs[i= dx], > >>>> so a flow can still accumulate 4096 packets of 1 MiB each and wrap t= he > >>>> counter mod 2^32. After a wrap, fq_codel_drop() sees a tiny maxbackl= og > >>>> and drops from an almost-empty flow, and the dequeue-side subtractio= ns > >>>> corrupt the counter further. > >>>> > >>>> Widen the fq_codel backlogs table, the fat-flow scan (maxbacklog/len= ) and > >>>> the drop threshold to u64. fq_codel is not lockless: every writer ru= ns > >>>> under the root qdisc lock, so plain u64 arithmetic keeps the WRITE_O= NCE > >>>> publish / READ_ONCE-consume pattern. The dump path > >>>> (fq_codel_dump_class_stats) stays a lockless stat-only read. > >>>> > >>>> CAKE accumulates the same generic qdisc_pkt_len(skb) into its per-fl= ow > >>>> b->backlogs[] and per-tin b->tin_backlog and consumes the values for > >>>> longest-flow pruning (cake_heapify/cake_heapify_up) and for the shap= er > >>>> staleness check, so it shares the bug. Widen those counters and the = heap > >>>> comparison locals to u64; the class/tin stats keep exporting the low= 32 > >>>> bits through the unchanged uAPI fields. > >>>> > >>>> Conditions to recreate the bug: CAP_NET_ADMIN in a user namespace; > >>>> CONFIG_NET_SCH_FQ_CODEL=3Dy. > >>>> > >>>> ip tuntap add tun0 mode tun > >>>> ip link set tun0 txqueuelen 32 up > >>>> ip addr add 10.99.0.1/24 dev tun0 > >>>> tc qdisc add dev tun0 root handle 1: stab overhead 2000000000 \ > >>>> fq_codel flows 1 limit 4200 ecn drop_batch 4096 > >>>> # hold the tun fd open without reading (IFF_BACKPRESSURE) so the qd= isc > >>>> # backlog persists, then send at least 4300 packets (the wrap start= s > >>>> # at 4096 resident; the over-limit drop that reads the wrapped > >>>> # threshold fires past the 4200 limit): backlogs[0] wraps at 4096 x > >>>> # 1 MiB and the fat-flow threshold reads the wrapped value. > >>>> > >>>> With a 1 MiB qdisc_pkt_len cap the counter wraps at 4096 resident > >>>> packets. At limit 4200 the first over-limit enqueue (the 4201st) see= s a > >>>> wrapped 105 MiB (half-backlog threshold 52 MiB, a ~52 packet drop bu= rst), > >>>> where the u64 counter sees 4201 MiB (threshold 2100 MiB, a ~2100 pac= ket > >>>> drop burst). > >>>> > >>>> Reported-by: Sashiko (nipa) > >>>> Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260818= 101130.16203-1-jhs@mojatatu.com > >>>> Link: https://lore.kernel.org/netdev/20260818101130.16203-1-jhs@moja= tatu.com/ > >>>> Tested-by: hybris > >>>> Signed-off-by: Jamal Hadi Salim > >>>> --- > >>> > >>> This makes no sense. > >>> > >> > >> As absurd as it looks that code is reachable ;-> > >> > >>> These qdisc have been developped to address bufferbloat issues. > >>> > >>> Storing 4GB in a qdisc is absolutely insane. > >>> > >> > >> The counter wrap is not because we stored 4GB, rather it is because > >> "tc .. stab overhead ..." inflates qdisc_pkt_len() for a 64B pkt to > >> 1MB. So ~4K packets (put in other words a few "real" KB) makes that > >> backlog[0] cross 2^32. > > > > Kill this stuff ? Who is still using this, for what reason ? > > tc-stab? Same as before, if you need want to model a remote bottleneck wi= th a local traffic shaper tc-stab is the generic solution (off the top of m= y head I only can enumerate cake as having its own traffic shaper that hand= les overheads). > One could argue that the "virtual" length should be accounted against cak= e's memlimit or fq-codel's memory_limit somehow... > In sane?/typical configurations overhead is expected to stay relatively s= mall, so this would not limit the actual queue size too much, while potenti= ally silencing this issue. linux qdisc are in the fast path, for nearly all packets sent over this pla= net. They aleady consume GW of energy. Modeling / network emulation should incur zero cost on these production grade qdisc. netem could be one answer, I do not know, or a special CONFIG_NET_SCHED_EXPENSIVE_EMULATION > > I might be off my rocker, in which has ignore (or preferably enlighten me= ).