Many ISPs need the kinds of quality shaping cake can do
 help / color / mirror / Atom feed
From: Frantisek Borsik <frantisek.borsik@gmail.com>
To: libreqos <libreqos@lists.bufferbloat.net>
Subject: [LibreQoS] Fwd: LibreQoS – QoS Systems for Disaster Recovery Support
Date: Thu, 1 Oct 2026 18:12:27 +0200	[thread overview]
Message-ID: <CAJUtOOjvXK1cTXUA8yRSp7KgqybUjTg5XtkYsKbMWGd3nU8LEw@mail.gmail.com> (raw)
In-Reply-To: <20261001150012.c5e5ed38bbd8f9f0@m.ghost.io>

Hello to all,

LibreQoS article is now available also online.

All the best,

Frank

Frantisek (Frank) Borsik


*In loving memory of Dave Täht: *1965-2025

https://libreqos.io/2025/04/01/in-loving-memory-of-dave/


https://www.linkedin.com/in/frantisekborsik

Signal, Telegram, WhatsApp: +421919416714

iMessage, mobile: +420775230885

Skype: casioa5302ca

frantisek.borsik@gmail.com


---------- Forwarded message ---------
From: ADMIN IT Infrastructure & Operations <
admin-it-infrastructure-operations@ghost.io>
Date: Thu, Oct 1, 2026 at 5:00 PM
Subject: LibreQoS – QoS Systems for Disaster Recovery Support
To: <frantisek.borsik@gmail.com>


Quality-of-service systems help prioritize scarce resources and keep
critical services running with efficient bandwidth management.
 ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏
 ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏
 ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏
 ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏
 ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏
 ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏  ͏
­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­
­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­
­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­
­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­
­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­ ­
­ ­ ­ ­ ­ ­ ­ ­ ­ ­

[image: ADMIN IT Infrastructure & Operations]
<https://www.admin-it.io/r/29e367d6?m=418155f0-5b01-4c85-b347-9682f63d540b>

ADMIN IT Infrastructure & Operations
<https://www.admin-it.io/r/37b221fc?m=418155f0-5b01-4c85-b347-9682f63d540b>
LibreQoS – QoS Systems for Disaster Recovery Support
<https://www.admin-it.io/r/42b6d34f?m=418155f0-5b01-4c85-b347-9682f63d540b>
By Frantisek Borsik & Tam Hanna • 1 Oct 2026 View in browser
<https://www.admin-it.io/r/ae2e706c?m=418155f0-5b01-4c85-b347-9682f63d540b>
View in browser
<https://www.admin-it.io/r/a9310254?m=418155f0-5b01-4c85-b347-9682f63d540b>
[image: Helping fellow mountain climber]

Lead Image © crazymedia, 123RF.com


Lifeline

After a catastrophic failure of the IT infrastructure – whether from damage
to undersea cables, weather, or pilfering by copper thieves – restoring
full bandwidth between the individual nodes of a network can take days or
even weeks. The quality of the user experience during these emergency
operations depends on how quickly applications and core services respond to
requests. Efficient bandwidth and latency management is crucial for a
company's disaster recovery (DR) strategy, and tools such as LibreQoS can
help you allocate scarce resources and prioritize latency-critical services.
QoS Data Stream Priorities

From a user experience perspective, some packets are not as important as
others. Whereas higher latency during a Windows update or when uploading an
object store BLOB is not particularly critical, a 100ms delay in voice or
video connections is disruptive.

As shown in Figure 1, quality-of-service (QoS) systems sit between local
and global networks. Packets passing through are classified and then
prioritized or throttled by applying various algorithms. Note that QoS
systems do not interact with intranet packets: Prioritizing the connection
between two stations located on the same network is not supported.
Figure 1: The QoS system typically sits between networks.

The procedures implemented in the prioritization algorithms vary from
vendor to vendor but can basically be broken down into two algorithm
families: common applications kept enhanced (CAKE), the more modern of the
two, and flow queue controlled delay (FQ-CoDel). For a better
understanding, I'll start with the controlled delay algorithm. As specified
in RFC 8289 [1]
<https://www.admin-it.io/r/a6d13373?m=418155f0-5b01-4c85-b347-9682f63d540b>,
the individual packet queues are analyzed first. If the wait time in one of
the queues is too long, it drops packets to reduce the wait. An intelligent
transmitter detects these packet losses and reduces the transmission rate.

FQ-CoDel, specified in RFC 8290 [2]
<https://www.admin-it.io/r/437be3c8?m=418155f0-5b01-4c85-b347-9682f63d540b>,
prioritizes short connections to enable faster processing and limit
infrastructure congestion caused by large packets or long-term
communication processes. CAKE, currently only documented in an eprint
article [3]
<https://www.admin-it.io/r/60001e7d?m=418155f0-5b01-4c85-b347-9682f63d540b>,
supplements the system with a traffic shaper and an improved hash
algorithm. In general, CAKE generates slightly higher CPU load and needs
significantly more RAM. In fact, LibreQoS estimates that it requires seven
times more RAM than FQ-CoDel. On the upside, a QoS system that uses this
approach will achieve lower latency.

In practical terms, the exact technical differences between the various QoS
methods are secondary. What matters most is the paradigm change from random
load distribution to some form of QoS.
DPI vs. Header Analysis

A QoS system must know the sensitivity of network packets to latency. Two
approaches rate this value: Deep packet inspection (DPI) analyzes the
content of the packets, whereas header-based methods only evaluate the
metadata, such as the TCP or IP fields.

DPI-based methods enable very granular control. For example, WhatsApp and
FaceTime calls can be distinguished and prioritized in different ways, but
you can encounter two drawbacks:

   - Legal considerations: In the EU, for example, DPI can be problematic
   because it potentially violates net neutrality regulations.
   - High system load: Analyzing packet content requires significantly more
   compute power and can overload a QoS system.

Header-based methods are simpler and more resource-efficient. They use
metadata, such as the priority bit stored in the IPv4 differentiated
services field (RFC 2474), or other TCP/IP header fields to assign packets
automatically to priority classes. CAKE uses the four classes listed in
Table 1.
Table 1: Priority Classes in CAKE
Priority Class Name Description
1 Voice VoIP, NTP, high-priority live data transmissions
2 Video Higher priority data
3 Best effort Most of the data to be transmitted
4 Bulk BitTorrent and similar low-priority traffic From Industry Giants to
Open Source

The QoS systems available on the market differ in terms of both the
algorithms and the licensing models they use. Two providers with a long
history are Sandvine, based in the US, and Preseem, founded in Canada by
former Sandvine employees. Both exclusively offer their software under
commercial licenses. Technically, the offerings differ in that Sandvine
relies on DPI, uses its own algorithm, and primarily caters to the needs of
Tier 1 ISPs. Preseem uses CAKE, but not DPI, because the product is aimed
at smaller ISPs.

Spain's Bequant relies on DPI and uses FQ-CoDel to prioritize packets. The
vendor has two unique selling points: First, it offers a price calculator
[4]
<https://www.admin-it.io/r/9bf930a6?m=418155f0-5b01-4c85-b347-9682f63d540b>
on its website that instantly gives a cost estimate on the basis of the IP
address. For a 200Mbps network, for example, a European customer can expect
to pay a one-off license fee of approximately EUR2,100. A North American
customer can get a monthly subscription license for about $114. The fee
increases parallel to the transfer speed. Second, Bequant licenses its
proprietary engine to third-party companies. Cambium Networks, based in
North America, for example, uses the technology for its own product.
Paraqum Technologies also plays a role in Asia. The Sri Lankan company
originated as a university spin-off and combines CAKE with DPI.

Libre-QoS [5]
<https://www.admin-it.io/r/ef10c390?m=418155f0-5b01-4c85-b347-9682f63d540b>
is the only open source system. It uses the CAKE algorithm, and the basic
version is available free of charge. Costs are only incurred if the
statistics module is used or the operator needs customer support. Although
support costs are negotiable, the statistics module is billed at $0.30 per
user and month for smaller deployments.
QOS During a Cable Outage

Now I'll look at real-world deployment scenarios for QoS. As an initial
practical example, assume that the fiber optic line connecting a company to
the Internet has failed. Experience shows that repairs to the fiber line
can take several weeks.

An ISP's first step would therefore typically be to set up a temporary
connection (e.g., microwave radio). A StarLink transceiver can ensure
connectivity, but the connection will have greater latency and lower
bandwidth.

At this point, services such as voice over Internet protocol (VoIP) or mail
are more important than, say, updating Visual Studio as quickly as
possible. Even on a network that normally operates without QoS, the limits
in this scenario make QoS a must-have. Strict resource management is
essential to restore a level of service acceptable to the majority of
users. As long as you have no extreme imbalance between the number of users
and the bandwidth (3-10Mbps per person is usually fine), deploying a QoS
system with the default settings will eliminate the problem. For this
purpose, it makes sense to have a dedicated server available that can be
inserted into the connection (Figure 1), if and when the need arises.

In situations with extremely low bandwidth, manual allocation to target
systems or services is the recommended approach. In the case of LibreQoS,
you can do this either in the graphical user interface or by generating a
CSV file with the settings. Practical experience shows that a manual
configuration is more sensible in disaster recovery situations, because
customer relationship management (CRM) and other systems are generally not
available until a later stage.

A CSV snippet example shown in the "LibreQoS Settings" box is taken from an
ISP's live configuration. Note that the first line contains information
about the field content. The Min and Max fields can then be used to specify
how much bandwidth needs to be available for the various systems, which
means you can prioritize the router for critical departments – outages
affecting the sales team are typically more critical than delays in
software updates.
LibreQoS Settings

Circuit ID,Circuit Name,Device ID,Device Name,Parent
Node,MAC,IPv4,IPv6,Download Min Mbps,Upload Min Mbps,Download Max
Mbps,Upload Max Mbps,Comment
1,"968 Circle St., Gurnee, IL 60031",1,Device 1,AP_A,,"100.64.0.1,
100.64.0.14",fdd7:b724:0:100::/56,1,1,155,20,
2,"31 Marconi Street, Lake In The Hills, IL 60156",2,Device
2,AP_A,,100.64.0.2,fdd7:b724:0:200::/56,0.5,0.5,2.5,1,
3,"255 NW. Newport Ave., Jamestown, NY 14701",3,Device
3,AP_9,,100.64.0.3,fdd7:b724:0:300::/56,1.25,0.75,10.5,5.25,
4,"8493 Campfire Street, Peabody, MA 01960",4,Device
4,AP_9,,100.64.0.4,fdd7:b724:0:400::/56,25.5,12.5,100.5,50.25,
2794,"6 Littleton Drive, Ringgold, GA 30736",5,Device
5,AP_11,,100.64.0.5,fdd7:b724:0:500::/56,1,1,105,18,
2794,"6 Littleton Drive, Ringgold, GA 30736",6,Device
6,AP_11,,100.64.0.6,fdd7:b724:0:600::/56,1,1,105,18,
5,"93 Oklahoma Ave., Parsippany, NJ 07054",7,Device
7,AP_1,,100.64.0.7,fdd7:b724:0:700::/56,2,2,25,10,
. . .

Defining the Recovery Sequence

At first glance, the targeted allocation of bandwidth to individual
consumers suggests that QoS could also be used to sequence the recovery
process. For example, if availability zone A is only given minimal
bandwidth, zone B could theoretically complete its recovery faster.

In practice, however, this approach proves to be of little benefit. The
recovery sequence for availability zones or logical functional blocks needs
to be defined at a higher level, because sequencing that is based on system
dependencies enables a clearer and more repeatable recovery.

Experience also shows that the data volumes required to restore a system
vary greatly and are difficult to estimate precisely up front. Disaster
recovery processes usually take place during periods of high instability.
Avoiding potentially undefined behavior and managing the recovery process
outside the QoS configuration explicitly is advised.
Anomaly Detection and AI-Based Analysis

In times of increasing threats from zero-day exploits and targeted attacks,
anomaly detection plays a crucial role. As in electronics or software
development, the "early detection reduces potential damage" principle
applies.

QoS systems often provide graphical overviews on the back end that
visualize the makeup of all the transmitted traffic (Figure 2). This
information can be used to identify suspicious patterns, such as the
exfiltration of large volumes of data to unusual destinations. For a simple
approach, you can evaluate these diagrams manually. Conspicuous patterns
(e.g., massive data transfers from the US to Russia) can indicate
exfiltration attempts or ransomware attacks. Unusual access to sensitive
systems is equally critical: For example, if the overview shows SSH access
from Vietnam to a European router, an attack is likely.
Figure 2: A LibreQoS AI analysis provides information for optimizing
network topology.

Providers of QoS systems are now going a step further: Large language
models (LLMs) are increasingly being integrated into the analysis to detect
anomalies automatically and provide actionable insights. At LibreQoS, an
LLM is currently in closed alpha testing; similar developments are also
expected from other vendors.

QoS systems not only support disaster recovery but also help with ongoing
optimization tasks, particularly for ISPs. Because many providers draw on a
broad database of production installations, their systems can automatically
detect common configuration glitches and generate optimization
recommendations. Don't forget that QoS systems are capable of proactively
identifying critical operational states, both through legacy anomaly
detection and by referencing the knowledge base. These indicators can
optionally trigger automatic alerts.
Preparing the LibreQoS Appliance

Given the benefits outlined here, it might make sense to prepare a LibreQoS
instance up front as part of your DR strategy. Compact appliance systems
are a very good choice. One option recommended by the developers is the
MinisForum MS-01: a compact mini-PC that offers two 2.5Gbps Ethernet ports
and two 10Gbps interfaces. Despite the high port density, resource
requirements are quite moderate: 4GB of RAM is all you need for
approximately 1,000 users.

In principle, LibreQoS can also be run on other hardware or a virtual
machine. In practical terms, though, you will run into limits, because the
software interacts very closely with the network interface cards (NICs).
Unsupported NICs (see the LibreQoS documentation [5]
<https://www.admin-it.io/r/ea132c73?m=418155f0-5b01-4c85-b347-9682f63d540b>)
can be used, but at your own risk and without support. Virtual machines
require NIC passthrough and significantly more compute power, which makes
administration more complex.
Configuring a Network Bridge

The LibreQoS host operating system is Ubuntu Server 24.04. After the
install, you need to set up a network bridge for the upstream and
downstream interfaces; NetPlan is used by default. The bridge structure is
detailed in the documentation. The configuration of the
/etc/netplan/libreqos.yaml file for interfaces named ens19 and ens20 might
look like Listing 1.
Listing 1: An Interface Config File

network:
ethernets:
  ens19:
   dhcp4: no
   dhcp6: no
  ens20:
   dhcp4: no
   dhcp6: no
bridges:
  br0:
   interfaces:
    - ens19
    - ens20
version: 2

Setting the dhcp* attributes to no ensures that the interfaces are
automatically active at system startup, even without an IP address
assignment. The commands

sudo chmod 600 /etc/netplan/ libreqos.yaml
sudo netplan apply

enable the configuration.
Installing and Configuring LibreQoS

A DEB package is available for the LibreQoS installation. To begin,
download, unpack, and install the package:

wget https://libreqos.io/wp-content/uploads/2025/03/libreqos_1.5-BETA10_amd64.zip
sudo apt-get install unzip
unzip libreqos_1.5-BETA10_amd64.zip
sudo apt install ./libreqos_1.5-BETA10_amd64.deb

Next, you need to edit the /etc/lqos.conf configuration file. The [bridge]
section is particularly important, because it defines the interfaces for
local and external communication. After saving the changes, activate the
system by restarting the system, or run the command

sudo systemctl restart lqosd

The maintenance overhead for a LibreQoS appliance is minimal. Besides the
standard updates to the host operating system, updates to the LibreQoS
software are also released regularly. From experience, a new version can be
expected approximately every five to seven months.
Conclusion

Disaster recovery requires a comprehensive strategy to restore systems and
services to a functional state as quickly as possible. In scenarios with
massive bandwidth restrictions, QoS systems play a crucial role by enabling
targeted prioritization of critical services, improving the user
experience, and helping to maintain productivity.
[1] RFC 8289 – CoDel: https://datatracker.ietf.org/doc/html/rfc8289
<https://www.admin-it.io/r/b8cf2504?m=418155f0-5b01-4c85-b347-9682f63d540b>
[2] RFC 8290 – flow queue CoDel: https://datatracker.ietf.org/doc/rfc8290/
<https://www.admin-it.io/r/c8af863f?m=418155f0-5b01-4c85-b347-9682f63d540b>
[3] Hoiland-Jorgensen, T., D. Täht, J. Morton. Piece of CAKE: A
comprehensive queue management solution for home gateways. 2018;
*arXiv:1804.07617*: https://arxiv.org/abs/1804.07617
<https://www.admin-it.io/r/bc754ef9?m=418155f0-5b01-4c85-b347-9682f63d540b>
[4] Bequant price calculator: https://www.bequant.com/pricing
<https://www.admin-it.io/r/ebfc72be?m=418155f0-5b01-4c85-b347-9682f63d540b>
[5] LibreQoS: https://libreqos.com
<https://www.admin-it.io/r/0a20c806?m=418155f0-5b01-4c85-b347-9682f63d540b>
[6] LibreQoS documentation: https://libreqos.readthedocs.io/en/latest/
<https://www.admin-it.io/r/f5a8c525?m=418155f0-5b01-4c85-b347-9682f63d540b>
[image: Comment]

Comment
<https://www.admin-it.io/r/38b89f7f?m=418155f0-5b01-4c85-b347-9682f63d540b>
ADMIN IT Infrastructure & Operations © 2026 – Unsubscribe
<https://www.admin-it.io/unsubscribe/?uuid=418155f0-5b01-4c85-b347-9682f63d540b&key=960c58e43eca66beab57aa713fc68d2ec316ff977f43b56fc025f05059e5e863&newsletter=665476b3-a799-4ce5-93d6-4a8c7bbe27c0>
[image: Powered by Ghost]
<https://ghost.org/?via=pbg-newsletter&ref=admin-it.io>

           reply	other threads:[~2026-10-01 15:59 UTC|newest]

Thread overview: expand[flat|nested]  mbox.gz  Atom feed
 [parent not found: <20261001150012.c5e5ed38bbd8f9f0@m.ghost.io>]

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

  List information: https://lists.bufferbloat.net/postorius/lists/libreqos.lists.bufferbloat.net/

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=CAJUtOOjvXK1cTXUA8yRSp7KgqybUjTg5XtkYsKbMWGd3nU8LEw@mail.gmail.com \
    --to=frantisek.borsik@gmail.com \
    --cc=libreqos@lists.bufferbloat.net \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox