From mboxrd@z Thu Jan 1 00:00:00 1970 Authentication-Results: mail.toke.dk; dkim=pass header.d=mojatatu.com header.i=@mojatatu.com header.a=rsa-sha256 header.s=google header.b="St/KMpMx"; arc=pass; dmarc=none Received: from mail-dl2-x10.google.com (mail-dl2-x10.google.com [IPv6:2607:f8b0:4864:38::10]) by mail.toke.dk (Postfix) with ESMTPS id A95801700DA2 for ; Sat, 26 Sep 2026 11:43:43 +0200 (CEST) Received: by mail-dl2-x10.google.com with SMTP id a92af1059eb24-142dd04edb5so2679740c88.2 for ; Sat, 26 Sep 2026 02:43:43 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1790415821; cv=none; d=google.com; s=arc-20260327; b=lJ2FiEOYQ8fXW89y35LyZIOrl80k1uW2yDkp7JQLJArV3fzpsFOdFTtpqPWwvjJu5l JZx7+0eX38Sc5raZLwvE5lUET4rVWmtG9SO7x9gx+ecTDBNddGaKRTdhmhVnNma61ZN4 jnjfmn7fK1Z2k7JcCxC8sMhXIFJgMRvmvJGABylv30A4RDeJ9dt8pMBmW9F04nD39FT+ yE+wH+axE03Ej7Z3ZBEaeUr5tPXSzqyUAk+DTkgSUgV19MEjyjYD4luZdF5Lyah080z+ GBNCVdAtHre1q2J06MHLylPk+28lIRY1mVSQRTmTKcLnYXr5cVUQjnyLDKqJb07e3qWT T5fg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=content-transfer-encoding:cc:to:subject:message-id:date:from :in-reply-to:references:mime-version:dkim-signature; bh=SzvWQLkWxwpiX7qLxPQTvx30y630CcfiVhKwC04gV5k=; fh=T1fAEc74VodC1Vty0D21IBEF34kbsLb8DIfEVwBdySU=; b=YLI97BRNQ2RnR/02MGd2ez48KPzCkID4vc3KL218Vq2m0K0KUjMRsLfW0gI0f5jmfx vO//CzxkY0MRWbGZ+eGDdYqWr5smlOBwD8TORtK/eqwd6Hs8jTkWPBrAO92c7oMYmCh/ Sgyaa6qay7OBvOe1hTOH7TX2YZSD18o/ThjEUqAZRMrqEzFghOTc9NEomYMhOKlWlXlC HFRsbG0x5P85tNMpseKbJ0Y+06lCy+YWsIQpSfgHAn0nMgLou2x4RwR7OSyVeSK7gBjQ r9rZM/rzAws/3QohQkKTZ92VOrJTxAQAad8GezZqFdcGCwQmaE6jUbHHeJnn89AxZCLJ dqLg==; darn=lists.bufferbloat.net ARC-Authentication-Results: i=1; mx.google.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=mojatatu.com; s=google; t=1790415821; x=1791020621; darn=lists.bufferbloat.net; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:from:to:cc:subject :date:message-id:reply-to:content-type; bh=SzvWQLkWxwpiX7qLxPQTvx30y630CcfiVhKwC04gV5k=; b=St/KMpMxFpTxfa74XWC9AxGsXWfmp1KBF35NHWSLX0ghHLFKBjTcdmZoU7XDqO4hK2 BXZyi8k/AIVTMgA5KLmTHZAdmt18u8u5lhod+8NzJgpozYuiSYVuysvybpoX3c9JpqjL 9YwEDMlhBMy/LCpF2HvqhXCtZ7uZYeikuGSxQ= X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790415821; x=1791020621; h=content-transfer-encoding:content-type:cc:to:subject:message-id :date:from:in-reply-to:references:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=SzvWQLkWxwpiX7qLxPQTvx30y630CcfiVhKwC04gV5k=; b=zlqbNKVEUbQvof7xa4PyMpCIcyg/2m3kgKraD+7MkpxBeS28cvafD7LkUCtJ78yYDF KiPN+ck3Lwbm3sqKw9IO5MoJdeZrunCP1loGib1bHF7VRO2fnZCc5gXjk0ZrrdKaSplM 5CD3DRhAQXWn3Vn3PgZGXyMa2/+gAJYmJm8fj5EQnIYOauDDK5Bsqo0v2VwhPXmIoa4i fNOt83fOcsohfa87bpVS3N4PbRbFXb9tBmkao0zwp02BbTnMTlhaH0AXZeBvrD5IPwic 7Kv7RhKUO0jFj3wvw/sz8nAQe3DmcnSRiZAhoDpXN1ddGvEodLbL9AK/rwsjXjjgr2+E Ka7w== X-Forwarded-Encrypted: i=1; AKwUvBwZZzYl3v8OmF3p6ZF+zMuuoIreniq8JkWCieVyUMvxjcHvsM8rdwBml0McskCkMtcgLYW6@lists.bufferbloat.net X-Gm-Message-State: AFuF++nf2/7k8MxIgr6OlJM6RPx9s/AbvkpcSX+FGU+WYDZs4im5ff28 Al8D3xsaK+b5fgKTCRR1TJHKgtU4I/op0QsTRJcfSzFCgGJMMPH/heKpZtKqKRf2kIFoJgI8Isx +asQPmj90jOrIdeIADB72qLc7DO50hgq/MMXVtPyt X-Gm-Gg: AYBFou1B2kQMthsP12IXDfNebiX7Q8ahLm/6irFeysspkMg3DFzXurwJrpuI9pcj19E ypy4bzDmsGbruzNbGvogfDem3FzWqVjajF1iZDVmJLow52vlcOmB6jzRxjY2EFw6C1QvYawTV1a abYEkK4mB//GdhrZM9CoJPPCnisGatz+rw4YNrBC3EWGacdsBFf7E61xPeXm9FS6xr2U938AOji ZUDi/Fdzznsuz0/8L/5mhQqHRNifU/WPFC1THIIemSKYx25q9je/xUUP+6V2NRE3yaCfETH4xRR n6PFGZw4wvvbN9md0bAR6E1jH73LUGCCtgggmcL6qhC3crCuZkFEXUzB85Q00C9uKhnTPGbBIRs esEC/7HYFaYIK66OGERAWbbyAyA7NiKP0zhuLcqNe2d+6Kb/kDwl6wqKccCEWGa36RfAvHkyY2B t0jBBkTURiNh3B+QWHYvs= X-Received: by 2002:a05:7022:ea8b:b0:146:5f73:d419 with SMTP id a92af1059eb24-146ce680e33mr2627989c88.14.1790415820581; Sat, 26 Sep 2026 02:43:40 -0700 (PDT) MIME-Version: 1.0 References: <68BE1514-828D-4184-84E4-90F2EB3F035D@gmx.de> <67EA1744-69D7-4A64-9117-D6A79A3A93E4@gmx.de> In-Reply-To: <67EA1744-69D7-4A64-9117-D6A79A3A93E4@gmx.de> From: Jamal Hadi Salim X-Gm-Features: AclHuK9nIZVvB4vxvUGaioGgnYIENDgkjJUY57GgKyB5XSUjBc0sRzpTlg5iQVs Message-ID: To: Sebastian Moeller Cc: Eric Dumazet , Jamal Hadi Salim , netdev@vger.kernel.org, =?UTF-8?B?VG9rZSBIw7hpbGFuZC1Kw7hyZ2Vuc2Vu?= , cake@lists.bufferbloat.net, Jiri Pirko , "David S . Miller" , Jakub Kicinski , Paolo Abeni , Simon Horman , Victor Nogueira , hybris , Sashiko Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable X-MailFrom: hadi@mojatatu.com X-Mailman-Rule-Hits: nonmember-moderation X-Mailman-Rule-Misses: dmarc-mitigation; no-senders; approved; loop; banned-address; emergency; member-moderation Message-ID-Hash: YJ2EQLHGKG3ZYNIWCLLGX5NLM75AGD7L X-Message-ID-Hash: YJ2EQLHGKG3ZYNIWCLLGX5NLM75AGD7L X-Mailman-Approved-At: Mon, 28 Sep 2026 13:06:36 +0200 X-Mailman-Version: 3.3.10 Precedence: list Subject: [Cake] Re: [PATCH net-next] net/sched: fq_codel, cake: widen backlogs to u64 List-Id: Cake - FQ_codel the next generation Archived-At: List-Archive: List-Help: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: Date: Sat, 26 Sep 2026 09:43:44 X-Original-Date: Sat, 26 Sep 2026 05:43:28 -0400 On Fri, Sep 25, 2026 at 10:03=E2=80=AFAM 'Sebastian Moeller' via Hyper-Yielding Back-end Review & Insight System wrote: > > Hi Eric, > > > > On Sep 25, 2026, at 15:43, Eric Dumazet wrote: > > > > On Fri, Sep 25, 2026 at 3:30=E2=80=AFPM Sebastian Moeller wrote: > >> > >> Hi Eric, > >> > >> > >>> On Sep 25, 2026, at 15:03, Eric Dumazet via Cake wrote: > >>> > >>> On Fri, Sep 25, 2026 at 1:28=E2=80=AFPM Jamal Hadi Salim wrote: > >>>> > >>>> On Fri, Sep 25, 2026 at 5:13=E2=80=AFAM Eric Dumazet wrote: > >>>>> > >>>>> On Fri, Sep 25, 2026 at 10:54=E2=80=AFAM Jamal Hadi Salim wrote: > >>>>>> > >>>>>> This is a follow-up to commit 8f735d64382d ("net/sched: bound > >>>>>> qdisc_pkt_len to prevent qdisc soft lockup"), which capped > >>>>>> qdisc_pkt_len() at QDISC_PKT_LEN_MAX (1 MiB). That cap bounds the = stab > >>>>>> amplifier but leaves the per-flow backlog counter u32: > >>>>>> fq_codel_enqueue() accumulates qdisc_pkt_len(skb) into q->backlogs= [idx], > >>>>>> so a flow can still accumulate 4096 packets of 1 MiB each and wrap= the > >>>>>> counter mod 2^32. After a wrap, fq_codel_drop() sees a tiny maxbac= klog > >>>>>> and drops from an almost-empty flow, and the dequeue-side subtract= ions > >>>>>> corrupt the counter further. > >>>>>> > >>>>>> Widen the fq_codel backlogs table, the fat-flow scan (maxbacklog/l= en) and > >>>>>> the drop threshold to u64. fq_codel is not lockless: every writer = runs > >>>>>> under the root qdisc lock, so plain u64 arithmetic keeps the WRITE= _ONCE > >>>>>> publish / READ_ONCE-consume pattern. The dump path > >>>>>> (fq_codel_dump_class_stats) stays a lockless stat-only read. > >>>>>> > >>>>>> CAKE accumulates the same generic qdisc_pkt_len(skb) into its per-= flow > >>>>>> b->backlogs[] and per-tin b->tin_backlog and consumes the values f= or > >>>>>> longest-flow pruning (cake_heapify/cake_heapify_up) and for the sh= aper > >>>>>> staleness check, so it shares the bug. Widen those counters and th= e heap > >>>>>> comparison locals to u64; the class/tin stats keep exporting the l= ow 32 > >>>>>> bits through the unchanged uAPI fields. > >>>>>> > >>>>>> Conditions to recreate the bug: CAP_NET_ADMIN in a user namespace; > >>>>>> CONFIG_NET_SCH_FQ_CODEL=3Dy. > >>>>>> > >>>>>> ip tuntap add tun0 mode tun > >>>>>> ip link set tun0 txqueuelen 32 up > >>>>>> ip addr add 10.99.0.1/24 dev tun0 > >>>>>> tc qdisc add dev tun0 root handle 1: stab overhead 2000000000 \ > >>>>>> fq_codel flows 1 limit 4200 ecn drop_batch 4096 > >>>>>> # hold the tun fd open without reading (IFF_BACKPRESSURE) so the q= disc > >>>>>> # backlog persists, then send at least 4300 packets (the wrap star= ts > >>>>>> # at 4096 resident; the over-limit drop that reads the wrapped > >>>>>> # threshold fires past the 4200 limit): backlogs[0] wraps at 4096 = x > >>>>>> # 1 MiB and the fat-flow threshold reads the wrapped value. > >>>>>> > >>>>>> With a 1 MiB qdisc_pkt_len cap the counter wraps at 4096 resident > >>>>>> packets. At limit 4200 the first over-limit enqueue (the 4201st) s= ees a > >>>>>> wrapped 105 MiB (half-backlog threshold 52 MiB, a ~52 packet drop = burst), > >>>>>> where the u64 counter sees 4201 MiB (threshold 2100 MiB, a ~2100 p= acket > >>>>>> drop burst). > >>>>>> > >>>>>> Reported-by: Sashiko (nipa) > >>>>>> Closes: https://netdev-ai.bots.linux.dev/sashiko/#/patchset/202608= 18101130.16203-1-jhs@mojatatu.com > >>>>>> Link: https://lore.kernel.org/netdev/20260818101130.16203-1-jhs@mo= jatatu.com/ > >>>>>> Tested-by: hybris > >>>>>> Signed-off-by: Jamal Hadi Salim > >>>>>> --- > >>>>> > >>>>> This makes no sense. > >>>>> > >>>> > >>>> As absurd as it looks that code is reachable ;-> > >>>> > >>>>> These qdisc have been developped to address bufferbloat issues. > >>>>> > >>>>> Storing 4GB in a qdisc is absolutely insane. > >>>>> > >>>> > >>>> The counter wrap is not because we stored 4GB, rather it is because > >>>> "tc .. stab overhead ..." inflates qdisc_pkt_len() for a 64B pkt to > >>>> 1MB. So ~4K packets (put in other words a few "real" KB) makes that > >>>> backlog[0] cross 2^32. > >>> > >>> Kill this stuff ? Who is still using this, for what reason ? > >> > >> tc-stab? Same as before, if you need want to model a remote bottleneck= with a local traffic shaper tc-stab is the generic solution (off the top o= f my head I only can enumerate cake as having its own traffic shaper that h= andles overheads). > >> One could argue that the "virtual" length should be accounted against = cake's memlimit or fq-codel's memory_limit somehow... > >> In sane?/typical configurations overhead is expected to stay relativel= y small, so this would not limit the actual queue size too much, while pote= ntially silencing this issue. > > > > linux qdisc are in the fast path, for nearly all packets sent over this= planet. > > Tip of the head to the linux network experts that enabled that and that t= ake care of it being efficient! > > > They aleady consume GW of energy. > > > > Modeling / network emulation should incur zero cost on these > > production grade qdisc. > > It is not modeling in a scientific sense (sorry I used inappropriate term= inology here), but simply the fact that to shape traffic to avoid filling r= emote queues one needs to be able to take the properties of that remote que= ue/interface/link-layer into account. > And that can be surprisingly close by. Last time I looked the kernel on s= ay eth0 added the 14 bytes of ethernet related overhead it handled itself t= o the packet size on top of the payload (no complaints that makes sense) bu= t that is insufficient to properly traffic shape e.g. the same ethernet int= erface... (as the relevant ethernet L1-frame overhead contains more bytes, = like the FCS, preamble and SFD). And that gets worse with stuff like DSL, c= able or fibre modems that are connected via ethernet to a linux router. > > > > > netem could be one answer, I do not know, or a special > > CONFIG_NET_SCHED_EXPENSIVE_EMULATION > > Well, tc-stab (or cake's overhead accounting) is already optional, so onl= y those incur the cost that actually use it (and only those users are affec= ted by the reported issue, it took tc-stab to cause the problem and arguabl= y in a configuration that is "insane"). > I might be trying to explain the internet here to people that make the in= ternet work, so apologies in advance but I want to make this as explicit as= I can. > TTraffic shaping without proper overhead accounting will not work as expe= cted (at least if the expectation is that the traffic shaper will honor the= set gross shaper rate). For a fixed packet size the lack over overhead acc= ounting can be papered over with reducing the shaper rate, but the amount o= f the required "over-shaping" depends on packet size (smaller packets requi= re more over-shaping), so generally that is sub optimal, because either the= shaper's guarantee is brittle or the over-shaping quite extreme (think the= header to payload ration of minimally- and MTU-sized packets). > > All I am saying is, we do still have use cases for tc-stab and proper ove= rhead accounting, at least on the leafs of the network like home internet l= inks. I am working on a v2 that drops at enqueue once the accounted backlog crosses (U32_MAX - QDISC_PKT_LEN_MAX), so no counter can wrap. That is Eric's original suggestion which preserves the point you raised. While looking at this closely for this update i noticed more qdiscs which for different reasons also suffer from ((u64)backlog + len <=3D limit idiom. So far it's clear from bfifo, gred and plug are the only ones safe because their limit is in bytes (as opposed to others which are packet counting) Eric: Unless i hear otherwise from you, the configurable byte ceiling you mentioned (a good noun seems to be CONFIG_NET_SCH_BACKLOG_CEILING) is a follow-up; cheers, jamal > > > >> > >> I might be off my rocker, in which has ignore (or preferably enlighten= me). > >