Cake - FQ_codel the next generation
 help / color / mirror / Atom feed
From: le berger des photons <thejoff@gmail.com>
To: cake@lists.bufferbloat.net
Subject: [Cake] Re: Cake Digest, Vol 130, Issue 7
Date: Tue, 21 Jul 2026 08:22:21 +0200	[thread overview]
Message-ID: <CAO-LeMwnoUds46p3eKsFBtCrN+p4kQYi3H9w-B+mqDeFhYYcAw@mail.gmail.com> (raw)
In-Reply-To: <178461405980.1649.1968000986901190552@gauss>

looking briefly at that it seems that cake is no longer properly named.  It
doesn't seem to me that it's a "piece of cake"  if I, who've been
networking for 20 years, has to invest a lot of time to even have a vague
idea of what you're talking about.

On Tue, Jul 21, 2026 at 8:07 AM <cake-request@lists.bufferbloat.net> wrote:

> Send Cake mailing list submissions to
>         cake@lists.bufferbloat.net
>
> To subscribe or unsubscribe via email, send a message with subject or
> body 'help' to
>         cake-request@lists.bufferbloat.net
>
> You can reach the person managing the list at
>         cake-owner@lists.bufferbloat.net
>
> When replying, please edit your Subject line so it is more specific
> than "Re: Contents of Cake digest..."
>
> Today's Topics:
>
>    1. Re: [PATCH net-next] net/sched: sch_cake: skip clearing unused tins
> during rate adjustment
>       (Toke Høiland-Jørgensen)
>    2. [PATCH net-next v2] net/sched: sch_cake: skip clearing unused tins
> during rate adjustment
>       (Jonas Köppeler)
>
>
> ----------------------------------------------------------------------
>
> Message: 1
> Date: Mon, 20 Jul 2026 22:30:09 +0200
> From: Toke Høiland-Jørgensen <toke@toke.dk>
> Subject: [Cake] Re: [PATCH net-next] net/sched: sch_cake: skip
>         clearing unused tins during rate adjustment
> To: Jonas Köppeler <j.koeppeler@tu-berlin.de>, Jamal Hadi Salim
>         <jhs@mojatatu.com>, Jiri Pirko <jiri@resnulli.us>, "David S.
> Miller"
>         <davem@davemloft.net>, Eric Dumazet <edumazet@google.com>, Jakub
>         Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>, Simon
>         Horman <horms@kernel.org>
> Cc: cake@lists.bufferbloat.net, netdev@vger.kernel.org,
>         linux-kernel@vger.kernel.org, Mike Pham <mikepham4321@gmail.com>
> Message-ID: <2E1DC220-4CC0-4F46-A042-6B09535EC373@toke.dk>
> Content-Type: text/plain; charset=utf-8
>
>
>
> On 18 July 2026 17.06.28 CEST, "Jonas Köppeler" <j.koeppeler@tu-berlin.de>
> wrote:
> >On 7/17/26 10:31, Toke Høiland-Jørgensen wrote:
> >> Jonas Köppeler <j.koeppeler@tu-berlin.de> writes:
> >>
> >>> When cake_configure_rates() is called from the dequeue path with
> >>> rate_adjust=true, it only needs to update the rate parameters. The
> >>> loop that clears the unused tins is both unnecessary and harmful in
> >>> this path:
> >>>
> >>>   - cake_clear_tin() overwrites q->cur_tin and q->cur_flow, which are
> >>>     actively used by cake_dequeue(), corrupting the dequeue state.
> >>>   - iterating over the unused tins and their internal queues to purge
> >>>     packets adds needless overhead to the hot path.
> >>>
> >>> Skip the entire loop when rate_adjust is set, as neither
> >>> cake_clear_tin() nor the mtu_time update are needed when only the
> >>> rate changes.
> >>>
> >>> Fixes: 15c2715a5264 ("net/sched: sch_cake: fixup cake_mq rate
> adjustment for diffserv config")
> >>> Signed-off-by: Jonas Köppeler <j.koeppeler@tu-berlin.de>
> >>> Tested-by: Mike Pham <mikepham4321@gmail.com>
> >>
> >> Do you have any performance numbers to show the impact of this?
> >Yes, the table below shows results from a test setup using vng with
> >2 network namespaces, with cake/cake_mq attached in one of them:
> >
> >    ns1 -> cake/cake_mq -> ns2
> >
> >- veth devices are configured with 8 rx/tx queues.
> >- cake/cake_mq is configured with a 2 Gbit rate limit.
> >- Running flent's rrul and tcp_nup tests with 32 TCP upstreams:
> >
> >legend: qdisc mq = cake_mq; mode be = besteffort, ds3 = diffserv3
> >        test nup = tcp_nup; base/load = idle/loaded RTT (ms); tput =
> Mbit/s
> >
> >+---------------------+-------+------+------+-------+-------+---------+
> >| kernel              | qdisc | mode | test |  base |  load |    tput |
> >+---------------------+-------+------+------+-------+-------+---------+
> >| net-next            | cake  | be   | rrul | 0.075 |  4.76 | 1473.69 |
> >| net-next            | cake  | be   | nup  | 0.078 |  6.23 | 1550.79 |
> >| net-next            | cake  | ds3  | rrul | 0.063 |  5.81 | 1526.75 |
> >| net-next            | cake  | ds3  | nup  | 0.046 |  6.09 | 1761.45 |
> >+---------------------+-------+------+------+-------+-------+---------+
> >| net-next            | mq    | be   | rrul | 0.810 | 11.78 | 1469.67 |
> >| net-next            | mq    | be   | nup  | 0.637 | 85.71 | 1243.15 |
> >| net-next            | mq    | ds3  | rrul | 0.397 | 15.28 | 1770.06 |
> >| net-next            | mq    | ds3  | nup  | 0.351 | 15.98 | 1799.39 |
> >+---------------------+-------+------+------+-------+-------+---------+
> >| this patch          | mq    | be   | rrul | 0.092 |  0.56 | 1873.40 |
> >| this patch          | mq    | be   | nup  | 0.109 |  1.82 | 1869.12 |
> >| this patch          | mq    | ds3  | rrul | 0.097 |  0.98 | 1866.10 |
> >| this patch          | mq    | ds3  | nup  | 0.101 |  0.51 | 1861.79 |
> >+---------------------+-------+------+------+-------+-------+---------+
> >| before 15c2715a5264 | mq    | be   | rrul | 0.073 |  0.30 | 1895.45 |
> >| before 15c2715a5264 | mq    | be   | nup  | 0.076 |  0.49 | 1905.57 |
> >| before 15c2715a5264 | mq    | ds3  | rrul | 0.069 |  0.31 | 1896.59 |
> >| before 15c2715a5264 | mq    | ds3  | nup  | 0.058 |  0.86 | 1884.01 |
> >+---------------------+-------+------+------+-------+-------+---------+
> >
> >Not only is p99 latency drastically reduced -- nearly matching
> >pre-15c2715a5264 results -- but on current upstream cake_mq,
> >throughput also increases as a cake mode uses more tins. This points
> >directly to cake_clear_tin() during reconfig as the cause, since it
> >clears (max_tins - cur_tins) tins each time. So the fewer tins the
> >current mode uses, the more get cleared on every reconfig.
> >
> >Mike ran also some test on OpenWrt, on an IPQ8074A with 4 rx/tx
> >queues, and saw similar trends. cake_mq is configured with a 2.2 Gbit
> >rate limit.
> >
> >Unfortunately, we only have data for 128 TCP upstreams on net-next,
> >and 64 TCP upstreams for 'this patch'.
> >
> >+---------------------+-------+------+------+---------+----------+
> >| kernel              | qdisc | mode | test |    load |     tput |
> >+---------------------+-------+------+------+---------+----------+
> >| net-next            | mq    | be   | nup  |  468.50 |    50.90 |
> >| net-next            | mq    | ds3  | nup  |  355.22 |    98.21 |
> >| net-next            | mq    | ds4  | nup  |  268.28 |   255.84 |
> >| net-next            | mq    | ds8  | nup  |    7.48 |  2023.66 |
> >+---------------------+-------+------+------+---------+----------+
> >| this patch          | mq    | be   | nup  |    4.24 |   944.35 |
> >| this patch          | mq    | ds3  | nup  |    4.27 |   937.75 |
> >| this patch          | mq    | ds4  | nup  |    4.24 |   936.97 |
> >| this patch          | mq    | ds8  | nup  |    4.32 |   927.89 |
> >+---------------------+-------+------+------+---------+----------+
> >
> >This again shows the same trend: throughput increases and latency
> >drops as cake_mq is configured with more tins. We're still looking
> >into why net-next+ds8 reaches close to 2 Gbit/s, while this patch
> >tops out around 928 Mbit/s.
> >
> >That said, this patch doesn't solve every issue yet, but it does
> >remove the regression introduced by commit 15c2715a5264
> >("net/sched: sch_cake: fixup cake_mq rate adjustment for diffserv
> >config").
> >
> >We're continuing to look into further improvements. Let us know if
> >you'd like to see additional tests :)
>
> Cool! Could you please respin the patch with this data in the commit
> message?
>
> Doesn't have to be all of it, but some indication of the benefit would be
> good to have on hand for future reference :)
>
> -Toke
>
> ------------------------------
>
> Message: 2
> Date: Mon, 20 Jul 2026 23:14:52 +0200
> From: Jonas Köppeler <j.koeppeler@tu-berlin.de>
> Subject: [Cake] [PATCH net-next v2] net/sched: sch_cake: skip clearing
>         unused tins during rate adjustment
> To: Toke Høiland-Jørgensen <toke@toke.dk>, Jamal Hadi Salim
>         <jhs@mojatatu.com>, Jiri Pirko <jiri@resnulli.us>, "David S.
> Miller"
>         <davem@davemloft.net>, Eric Dumazet <edumazet@google.com>, Jakub
>         Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>, Simon
>         Horman <horms@kernel.org>
> Cc: <cake@lists.bufferbloat.net>, <netdev@vger.kernel.org>,
>         <linux-kernel@vger.kernel.org>, Jonas Köppeler
>         <j.koeppeler@tu-berlin.de>, Mike Pham <mikepham4321@gmail.com>
> Message-ID:
>         <
> 20260720-sch_cake-skip-clearing-tins-v2-1-e6a8b0275c73@tu-berlin.de>
> Content-Type: text/plain; charset="utf-8"
>
> When cake_configure_rates() is called from the dequeue path with
> rate_adjust=true, it only needs to update the rate parameters. The
> loop that clears the unused tins is both unnecessary and harmful in
> this path:
>
>  - cake_clear_tin() overwrites q->cur_tin and q->cur_flow, which are
>    actively used by cake_dequeue(), corrupting the dequeue state.
>  - iterating over the unused tins and their internal queues to purge
>    packets adds needless overhead to the hot path.
>
> Skip the entire loop when rate_adjust is set, as neither
> cake_clear_tin() nor the mtu_time update are needed when only the
> rate changes.
>
> The clearing loop runs on every rate adjustment from the dequeue path,
> clearing (max_tins - cur_tins) tins each time, so the cost grows the
> fewer tins the configured mode actually uses. Testing cake_mq over veth
> (8 rx/tx queues, 2 Gbit limit) with flent's [1] rrul and tcp_nup tests and
> 32 TCP upstreams shows a large drop in loaded latency and a throughput
> gain, restoring behaviour to pre-15c2715a5264 levels:
>
>   +------------+------+------+-------+-------+---------+
>   | kernel     | mode | test |  base |  load |    tput |
>   |            |      |      |  (ms) |  (ms) |  (Mbit) |
>   +------------+------+------+-------+-------+---------+
>   | net-next   | be   | rrul | 0.810 | 11.78 | 1469.67 |
>   | net-next   | be   | nup  | 0.637 | 85.71 | 1243.15 |
>   | net-next   | ds3  | rrul | 0.397 | 15.28 | 1770.06 |
>   | net-next   | ds3  | nup  | 0.351 | 15.98 | 1799.39 |
>   +------------+------+------+-------+-------+---------+
>   | patched    | be   | rrul | 0.092 |  0.56 | 1873.40 |
>   | patched    | be   | nup  | 0.109 |  1.82 | 1869.12 |
>   | patched    | ds3  | rrul | 0.097 |  0.98 | 1866.10 |
>   | patched    | ds3  | nup  | 0.101 |  0.51 | 1861.79 |
>   +------------+------+------+-------+-------+---------+
>
> The same trend holds on real hardware (IPQ8074A, 4 rx/tx queues,
> OpenWrt): in besteffort mode the tcp_nup loaded latency drops from
> ~470 ms to ~4 ms.
>
> [1] https://flent.org
>
> Fixes: 15c2715a5264 ("net/sched: sch_cake: fixup cake_mq rate adjustment
> for diffserv config")
> Signed-off-by: Jonas Köppeler <j.koeppeler@tu-berlin.de>
> Tested-by: Mike Pham <mikepham4321@gmail.com>
> ---
> Changes in v2:
> - added performance data to commit message, no code changes
> - Link to v1:
> https://patch.msgid.link/20260716-sch_cake-skip-clearing-tins-v1-1-d9787df20c28@tu-berlin.de
> ---
>  net/sched/sch_cake.c | 8 +++++---
>  1 file changed, 5 insertions(+), 3 deletions(-)
>
> diff --git a/net/sched/sch_cake.c b/net/sched/sch_cake.c
> index f78f8e950776..845e1c017714 100644
> --- a/net/sched/sch_cake.c
> +++ b/net/sched/sch_cake.c
> @@ -2609,9 +2609,11 @@ static void cake_configure_rates(struct Qdisc *sch,
> u64 rate, bool rate_adjust)
>                 break;
>         }
>
> -       for (c = qd->tin_cnt; c < CAKE_MAX_TINS; c++) {
> -               cake_clear_tin(sch, c);
> -               qd->tins[c].cparams.mtu_time =
> qd->tins[ft].cparams.mtu_time;
> +       if (!rate_adjust) {
> +               for (c = qd->tin_cnt; c < CAKE_MAX_TINS; c++) {
> +                       cake_clear_tin(sch, c);
> +                       qd->tins[c].cparams.mtu_time =
> qd->tins[ft].cparams.mtu_time;
> +               }
>         }
>
>         qd->rate_ns   = qd->tins[ft].tin_rate_ns;
>
> ---
> base-commit: f6f3b36c15ed44de1fbb44e645e4fae8c4a4453e
> change-id: 20260716-sch_cake-skip-clearing-tins-856812586cde
>
> Best regards,
> --
> Jonas Köppeler <j.koeppeler@tu-berlin.de>
>
>
> ------------------------------
>
> Subject: Digest Footer
>
> _______________________________________________
> Cake mailing list -- cake@lists.bufferbloat.net
> To unsubscribe send an email to cake-leave@lists.bufferbloat.net
>
>
> ------------------------------
>
> End of Cake Digest, Vol 130, Issue 7
> ************************************
>

       reply	other threads:[~2026-07-21  6:22 UTC|newest]

Thread overview: 3+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <178461405980.1649.1968000986901190552@gauss>
2026-07-21  6:22 ` le berger des photons [this message]
2026-07-21  7:08   ` [Cake] Re: Cake Digest, Vol 130, Issue 7 Sebastian Moeller
2026-07-21  7:09   ` Frantisek Borsik

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

  List information: https://lists.bufferbloat.net/postorius/lists/cake.lists.bufferbloat.net/

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=CAO-LeMwnoUds46p3eKsFBtCrN+p4kQYi3H9w-B+mqDeFhYYcAw@mail.gmail.com \
    --to=thejoff@gmail.com \
    --cc=cake@lists.bufferbloat.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox