From: "Jonas Köppeler" <j.koeppeler@tu-berlin.de>
To: "Toke Høiland-Jørgensen" <toke@toke.dk>,
"Jamal Hadi Salim" <jhs@mojatatu.com>,
"Jiri Pirko" <jiri@resnulli.us>,
"David S. Miller" <davem@davemloft.net>,
"Eric Dumazet" <edumazet@google.com>,
"Jakub Kicinski" <kuba@kernel.org>,
"Paolo Abeni" <pabeni@redhat.com>,
"Simon Horman" <horms@kernel.org>
Cc: <cake@lists.bufferbloat.net>, <netdev@vger.kernel.org>,
<linux-kernel@vger.kernel.org>,
Mike Pham <mikepham4321@gmail.com>
Subject: [Cake] Re: [PATCH net-next] net/sched: sch_cake: skip clearing unused tins during rate adjustment
Date: Sat, 18 Jul 2026 17:06:28 +0200 [thread overview]
Message-ID: <431c1b00-1fd3-41e1-8e8e-9cc1738ae387@tu-berlin.de> (raw)
In-Reply-To: <87wluuj9ue.fsf@toke.dk>
On 7/17/26 10:31, Toke Høiland-Jørgensen wrote:
> Jonas Köppeler <j.koeppeler@tu-berlin.de> writes:
>
>> When cake_configure_rates() is called from the dequeue path with
>> rate_adjust=true, it only needs to update the rate parameters. The
>> loop that clears the unused tins is both unnecessary and harmful in
>> this path:
>>
>> - cake_clear_tin() overwrites q->cur_tin and q->cur_flow, which are
>> actively used by cake_dequeue(), corrupting the dequeue state.
>> - iterating over the unused tins and their internal queues to purge
>> packets adds needless overhead to the hot path.
>>
>> Skip the entire loop when rate_adjust is set, as neither
>> cake_clear_tin() nor the mtu_time update are needed when only the
>> rate changes.
>>
>> Fixes: 15c2715a5264 ("net/sched: sch_cake: fixup cake_mq rate adjustment for diffserv config")
>> Signed-off-by: Jonas Köppeler <j.koeppeler@tu-berlin.de>
>> Tested-by: Mike Pham <mikepham4321@gmail.com>
>
> Do you have any performance numbers to show the impact of this?
Yes, the table below shows results from a test setup using vng with
2 network namespaces, with cake/cake_mq attached in one of them:
ns1 -> cake/cake_mq -> ns2
- veth devices are configured with 8 rx/tx queues.
- cake/cake_mq is configured with a 2 Gbit rate limit.
- Running flent's rrul and tcp_nup tests with 32 TCP upstreams:
legend: qdisc mq = cake_mq; mode be = besteffort, ds3 = diffserv3
test nup = tcp_nup; base/load = idle/loaded RTT (ms); tput = Mbit/s
+---------------------+-------+------+------+-------+-------+---------+
| kernel | qdisc | mode | test | base | load | tput |
+---------------------+-------+------+------+-------+-------+---------+
| net-next | cake | be | rrul | 0.075 | 4.76 | 1473.69 |
| net-next | cake | be | nup | 0.078 | 6.23 | 1550.79 |
| net-next | cake | ds3 | rrul | 0.063 | 5.81 | 1526.75 |
| net-next | cake | ds3 | nup | 0.046 | 6.09 | 1761.45 |
+---------------------+-------+------+------+-------+-------+---------+
| net-next | mq | be | rrul | 0.810 | 11.78 | 1469.67 |
| net-next | mq | be | nup | 0.637 | 85.71 | 1243.15 |
| net-next | mq | ds3 | rrul | 0.397 | 15.28 | 1770.06 |
| net-next | mq | ds3 | nup | 0.351 | 15.98 | 1799.39 |
+---------------------+-------+------+------+-------+-------+---------+
| this patch | mq | be | rrul | 0.092 | 0.56 | 1873.40 |
| this patch | mq | be | nup | 0.109 | 1.82 | 1869.12 |
| this patch | mq | ds3 | rrul | 0.097 | 0.98 | 1866.10 |
| this patch | mq | ds3 | nup | 0.101 | 0.51 | 1861.79 |
+---------------------+-------+------+------+-------+-------+---------+
| before 15c2715a5264 | mq | be | rrul | 0.073 | 0.30 | 1895.45 |
| before 15c2715a5264 | mq | be | nup | 0.076 | 0.49 | 1905.57 |
| before 15c2715a5264 | mq | ds3 | rrul | 0.069 | 0.31 | 1896.59 |
| before 15c2715a5264 | mq | ds3 | nup | 0.058 | 0.86 | 1884.01 |
+---------------------+-------+------+------+-------+-------+---------+
Not only is p99 latency drastically reduced -- nearly matching
pre-15c2715a5264 results -- but on current upstream cake_mq,
throughput also increases as a cake mode uses more tins. This points
directly to cake_clear_tin() during reconfig as the cause, since it
clears (max_tins - cur_tins) tins each time. So the fewer tins the
current mode uses, the more get cleared on every reconfig.
Mike ran also some test on OpenWrt, on an IPQ8074A with 4 rx/tx
queues, and saw similar trends. cake_mq is configured with a 2.2 Gbit
rate limit.
Unfortunately, we only have data for 128 TCP upstreams on net-next,
and 64 TCP upstreams for 'this patch'.
+---------------------+-------+------+------+---------+----------+
| kernel | qdisc | mode | test | load | tput |
+---------------------+-------+------+------+---------+----------+
| net-next | mq | be | nup | 468.50 | 50.90 |
| net-next | mq | ds3 | nup | 355.22 | 98.21 |
| net-next | mq | ds4 | nup | 268.28 | 255.84 |
| net-next | mq | ds8 | nup | 7.48 | 2023.66 |
+---------------------+-------+------+------+---------+----------+
| this patch | mq | be | nup | 4.24 | 944.35 |
| this patch | mq | ds3 | nup | 4.27 | 937.75 |
| this patch | mq | ds4 | nup | 4.24 | 936.97 |
| this patch | mq | ds8 | nup | 4.32 | 927.89 |
+---------------------+-------+------+------+---------+----------+
This again shows the same trend: throughput increases and latency
drops as cake_mq is configured with more tins. We're still looking
into why net-next+ds8 reaches close to 2 Gbit/s, while this patch
tops out around 928 Mbit/s.
That said, this patch doesn't solve every issue yet, but it does
remove the regression introduced by commit 15c2715a5264
("net/sched: sch_cake: fixup cake_mq rate adjustment for diffserv
config").
We're continuing to look into further improvements. Let us know if
you'd like to see additional tests :)
- Jonas
>
> -Toke
next prev parent reply other threads:[~2026-07-18 15:06 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-16 20:59 [Cake] [PATCH net-next] net/sched: sch_cake: skip clearing unused tins during rate adjustment Jonas Köppeler
2026-07-17 8:31 ` [Cake] " Toke Høiland-Jørgensen
2026-07-18 15:06 ` Jonas Köppeler [this message]
2026-07-20 20:30 ` Toke Høiland-Jørgensen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
List information: https://lists.bufferbloat.net/postorius/lists/cake.lists.bufferbloat.net/
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=431c1b00-1fd3-41e1-8e8e-9cc1738ae387@tu-berlin.de \
--to=j.koeppeler@tu-berlin.de \
--cc=cake@lists.bufferbloat.net \
--cc=davem@davemloft.net \
--cc=edumazet@google.com \
--cc=horms@kernel.org \
--cc=jhs@mojatatu.com \
--cc=jiri@resnulli.us \
--cc=kuba@kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mikepham4321@gmail.com \
--cc=netdev@vger.kernel.org \
--cc=pabeni@redhat.com \
--cc=toke@toke.dk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox