In this article, we explore what network quality of service (QoS) means, preview an upcoming native capability that's addressing it declaratively, and then step through a practical solution you can deploy today using the Linux traffic control (tc) command. We'll set up a full reproducer lab, run it, and look at real test results that demonstrate production traffic maintaining throughput while background traffic is deprioritised under contention.
Suppose you have a use case like one of the following:
- Backups: Virtual machines (VM) running on Red Hat OpenShift back up to an external service during business hours and saturate a shared link.
- Log shipping: High volume log forwarders push gigabytes of telemetry to a central aggregator, competing with application traffic on the same interface.
- Bulk data transfers: Moving large volumes of data across the network in bursts.
In most cases, background traffic is easily identifiable by a destination IP or CIDR range, making it straightforward to target specific endpoints where you want to manage resource consumption based on available network capacity.
Historically, however, Kubernetes lacked built-in mechanisms to define traffic priority. The problem with traditional bandwidth limits comes down to how network traffic works. Network bandwidth is a compressible resource. Throttling it simply slows traffic down rather than stopping or crashing the workload. Non-compressible resources like memory are rigid: A container claims its allocation up front, cannot dynamically flex past its limit when demand spikes, and gets OOM killed if it tries.
Network bandwidth, on the other hand, is fluid, so it can safely scale up and down as traffic flows. Because network bandwidth is compressible, setting a hard limit or flat cap is almost always the wrong thing to do. A flat cap penalizes background tasks even when the network has available bandwidth, wasting capacity. Instead, we should take advantage of this compressibility to dynamically balance resource usage.
The real goal in our scenario is to prioritize critical production traffic while giving background workloads a guaranteed minimum so they never starve, while allowing them to consume spare bandwidth when production is idle. Achieving this requires a full traffic shaping hierarchy rather than a flat bandwidth cap.
The future: NetworkQoS in OVN-Kubernetes
To handle traffic shaping effectively, the logic must sit directly within the Container Network Interface (CNI) plug-in. Because packet queuing and filtering happen at the host level, the CNI is the natural place to enforce traffic shaping before packets hit the wire.
Network QoS is not yet a standard feature provided natively out of the box by core Kubernetes. Instead, traffic shaping capabilities depend heavily on the underlying network plug-in implementation.
OVN Kubernetes (OVNk) serves as the default CNI for OpenShift Container Platform (OCP) and the community is actively working on bringing a native, declarative model directly into OVNk. The NetworkQoS enhancement proposal (OKEP-4380) introduces a new NetworkQoS custom resource definition (CRD) under k8s.ovn.org/v1alpha1 that provides fine grained QoS controls for pod egress traffic.
Here's an example of a NetworkQoS resource:
apiVersion: k8s.ovn.org/v1alpha1
kind: NetworkQoS
metadata:
name: deprioritise-background
namespace: my-app
spec:
podSelector:
matchLabels:
app: my-vm
priority: 50
egress:
- dscp: 10
bandwidth:
rate: 2000000 # kbps
burst: 4000000 # kbps
classifier:
to:
- ipBlock:
cidr: 10.200.0.0/24The feature is currently tracked under OCPSTRAT-3266.
What you can do today with Linux traffic control
While NetworkQoS makes its way through the development process, Linux traffic control (the tc command) is an established kernel subsystem, and it is the foundation that many higher level network control tools build upon. It provides the general framework for traffic shaping directly at the host network layer.
tc operates on 3 concepts:
- Queueing disciplines (qdiscs): Algorithms decide how packets are enqueued and dequeued on a network interface.
- Classes: Subdivisions within a classful qdisc. Each class has a guaranteed rate and a ceiling it can burst to.
- Filters: Rules that classify packets into classes based on headers, marks, or other properties.
Example tc approach
Within traffic control (tc), a hierarchical token bucket (HTB) is a classful queueing discipline used to manage and guarantee network bandwidth across different types of traffic. It works by organizing network queues into a tree hierarchy, allowing you to set minimum guaranteed bandwidths and maximum bandwidth limits for specific classes while letting idle bandwidth be shared dynamically among sibling queues.
For our use case, we want a hierarchy that:
- Guarantees production traffic a large share of bandwidth (90%)
- Guarantees background traffic a minimum floor (10%) so it never starves
- Allows either class burst to the full link capacity when the other is idle
- Keeps latency low within each class
Figure 1 illustrates the qdisc tree we will build on a 25 Gbps bonded interface:
Traffic control tutorial
To follow along, you must have these prerequisites:
- An OpenShift cluster with at least 2 worker nodes
- A localnet secondary ClusterUserDefinedNetwork (cUDN) configured
- The
occommand, authenticated to the cluster - 4 available IPs on the localnet subnet
The architecture used in this experiment is illustrated in figure 2, and can be summarised as:
2x Physical NICs → LACP Bond (bond1) → OVS Bridge → localnet cUDN
In this test scenario, 2 iperf3 server pods run on Node B. 2 iperf3 client pods run on Node A, where the tc rules are applied to bond1. All pods attach to a localnet secondary cUDN with static IPs. The tc filter matches the background (deprio) server's IP (192.168.0.79/32) and routes that traffic into the deprioritised class.
Step 1: Set variables
First, export the required environment variables:
export NODE_A=node1.example.com
export NODE_B=node2.example.com
export CUDN_NAME=example-vlan
export PROD_SERVER_IP=192.168.0.78
export DEPRIO_SERVER_IP=192.168.0.79
export PROD_CLIENT_IP=192.168.0.80
export DEPRIO_CLIENT_IP=192.168.0.81Step 2: Create the test namespace
Create a namespace for this experiment:
oc new-project tc-qos-test
oc label namespace tc-qos-test "cudn/example-vlan=enabled"The label must match your cUDN's namespaceSelector so that the NetworkAttachmentDefinition is created automatically.
Step 3: Verify the NetworkAttachmentDefinition
A NetworkAttachmentDefinition was created automatically, using the environment variables you set, when you created a namespace. Verify that the NetworkAttachmentDefinition exists with the expected name:
$ oc get net-attach-def -n tc-qos-test
NAME AGE
example-vlan 22hStep 4: Deploy iperf3 server pods on node B
Deploy the iperf3 server pods on NODE_B (an environment variable you defined in step 1):
cat <<EOF | oc apply -f -
apiVersion: v1
kind: Pod
metadata:
name: iperf3-prod-server
namespace: tc-qos-test
annotations:
k8s.v1.cni.cncf.io/networks: '[{"name":"${CUDN_NAME}","ips":["${PROD_SERVER_IP}/23"]}]'
spec:
nodeName: ${NODE_B}
containers:
- name: iperf3
image: quay.io/jforce/network-tools:v1.0
command: ["iperf3", "-s"]
ports:
- containerPort: 5201
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "1"
memory: 2Gi
---
apiVersion: v1
kind: Pod
metadata:
name: iperf3-deprio-server
namespace: tc-qos-test
annotations:
k8s.v1.cni.cncf.io/networks: '[{"name":"${CUDN_NAME}","ips":["${DEPRIO_SERVER_IP}/23"]}]'
spec:
nodeName: ${NODE_B}
containers:
- name: iperf3
image: quay.io/jforce/network-tools:v1.0
command: ["iperf3", "-s"]
ports:
- containerPort: 5201
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "1"
memory: 2Gi
EOFStep 5: Deploy iperf3 client pods on node A
Deploy the iperf3 client pods on NODE_A (an environment variable you defined in step 1):
cat <<EOF | oc apply -f -
apiVersion: v1
kind: Pod
metadata:
name: iperf3-prod-client
namespace: tc-qos-test
annotations:
k8s.v1.cni.cncf.io/networks: '[{"name":"${CUDN_NAME}","ips":["${PROD_CLIENT_IP}/23"]}]'
spec:
nodeName: ${NODE_A}
containers:
- name: iperf3
image: quay.io/jforce/network-tools:v1.0
command: ["sleep", "infinity"]
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "1"
memory: 2Gi
---
apiVersion: v1
kind: Pod
metadata:
name: iperf3-deprio-client
namespace: tc-qos-test
annotations:
k8s.v1.cni.cncf.io/networks: '[{"name":"${CUDN_NAME}","ips":["${DEPRIO_CLIENT_IP}/23"]}]'
spec:
nodeName: ${NODE_A}
containers:
- name: iperf3
image: quay.io/jforce/network-tools:v1.0
command: ["sleep", "infinity"]
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "1"
memory: 2Gi
EOFStep 6: Verify secondary IPs
Confirm that each pod shows a secondary interface (net1) with its assigned IP:
for pod in iperf3-prod-server iperf3-deprio-server iperf3-prod-client iperf3-deprio-client; do
echo "--- ${pod} ---"
oc exec ${pod} -n tc-qos-test ip -brief addr
doneStep 7: Apply traffic control rules on node A
Use oc to launch a debug pod:
oc debug node/${NODE_A}
chroot /hostApply the HTB hierarchy tc configuration:
tc qdisc replace dev bond1 root handle 1: htb default 10
tc class replace dev bond1 parent 1: classid 1:1 htb rate 23750mbit ceil 23750mbit
tc class replace dev bond1 parent 1:1 classid 1:10 htb rate 21590mbit ceil 23750mbit quantum 90000
tc class replace dev bond1 parent 1:1 classid 1:20 htb rate 2160mbit ceil 23750mbit quantum 9000
tc qdisc replace dev bond1 parent 1:10 fq_codel
tc qdisc replace dev bond1 parent 1:20 fq_codel
tc filter replace dev bond1 parent 1: protocol 802.1Q prio 1 flower \
vlan_ethtype ipv4 dst_ip ${DEPRIO_SERVER_IP}/32 flowid 1:20Verify the configuration using the tc command:
tc -s qdisc show dev bond1
tc -s class show dev bond1
tc filter show dev bond1Step 8: Run the tests
Now we run tests to establish some baselines.
Test A: Background traffic only
The first test we run establishes a baseline for just the background traffic with no contention:
oc exec iperf3-deprio-client -n tc-qos-test -- \
iperf3 -c ${DEPRIO_SERVER_IP} -B ${DEPRIO_CLIENT_IP} -t 30This confirms there is no hard cap when production is silent, background traffic uses the full link.
Test B: Production traffic only
The next test establishes a baseline for just production traffic:
oc exec iperf3-prod-client -n tc-qos-test -- \
iperf3 -c ${PROD_SERVER_IP} -B ${PROD_CLIENT_IP} -t 30Test C: Contention with both streams simultaneously
The critical test tests whether production maintains high throughput while deprioritised is throttled:
oc exec iperf3-prod-client -n tc-qos-test -- \
iperf3 -c ${PROD_SERVER_IP} -B ${PROD_CLIENT_IP} -t 30 &
oc exec iperf3-deprio-client -n tc-qos-test -- \
iperf3 -c ${DEPRIO_SERVER_IP} -B ${DEPRIO_CLIENT_IP} -t 30 &
waitStep 9: Review the test results
Here are the actual results from our lab with 25 Gbps bonded NICs:
Test A: Background (deprio) only (no contention)
- Duration: 30 seconds
- Transfer: 71.3 GBytes
- Bitrate: 20.4 Gbits/sec
Background traffic ran at near-line-rate. No hard cap was applied, so this is exactly what we want when there is no contention.
Test B: Production only (no contention)
- Duration: 30 seconds
- Transfer: 71.1 GBytes
- Bitrate: 20.4 Gbits/sec
The production baseline is also at near-line-rate, as expected.
Test C: Contention (both streams simultaneously)
- Production
- Transfer: 68.2 GBytes
- Bitrate: 19.5 Gbits/sec
- Background
- Transfer: 9.42 GBytes
- Bitrate: 2.70 Gbits/sec
- Combined
- Transfer: 77.6 GBytes
- Bitrate: ~22.2 Gbits/sec
This is the result that matters. Under contention:
- Production: Stayed at 19.5 Gbps, only a marginal drop from its uncontested 20.4 Gbps baseline
- Background: Throttled to 2.7 Gbps just above its guaranteed floor of 2.16 Gbps
- Combined: Throughput reached ~22.2 Gbps, close to the 23.75 Gbps ceiling, meaning the link is well utilised.
The traffic control statistics from the node confirm the traffic split:
qdisc htb 1: root Sent 34.36 GB 22.7M pkt
fq_codel parent 1:10 Sent 30.34 GB 20.0M pkt (production)
fq_codel parent 1:20 Sent 4.02 GB 2.6M pkt (background)Production received approximately 7.5x the bytes of background traffic consistent with the 10:1 priority ratio, with some borrowing from the background class.
Step 10: Observe live traffic control statistics
While running the contention test, you can watch the counters update in real time from a separate terminal:
oc debug node/${NODE_A}
chroot /host
watch -n 1 tc -s qdisc show dev bond1Step 11: Clean up
After you've finished running the test scenarios, you can safely delete the project:
oc delete project tc-qos-testRemove the traffic control rules from the node:
oc debug node/${NODE_A}
chroot /host
tc qdisc del dev bond1 rootConclusion
Network QoS in Kubernetes is evolving. The upcoming NetworkQoS CRD in OVN Kubernetes will bring declarative, Kubernetes native traffic management directly into the platform, including DSCP marking, bandwidth policing, and metadata-based traffic classification using pod labels and destination CIDRs. When it lands in OpenShift, administrators will be able to express QoS policies the same way they express network policies today.
Until then, Linux craffic control remains a powerful and proven tool for those looking to take matters into their own hands. As we have shown, a well crafted HTB hierarchy on the node physical interface can deliver meaningful traffic prioritisation right now. However, applying tc commands manually in a lab environment is inherently temporary, these commands do not survive reboots or interface resets.
If you want to operationalise a self-managed workaround on OpenShift today, then you might try something along the lines of a custom script for your tc rules, wrapped in a systemd service to handle boot and interface events, and deployed to your worker nodes declaratively using a MachineConfig object.
To learn more, check out these resources: