GPU demand for AI workloads is surging, and sharing GPUs fairly across teams remains one of the hardest problems in cluster management. The Kubernetes device plug-in model was never built for this. It treats GPUs as anonymous, countable integers—for example, NVIDIA GPUs are requested as nvidia.com/gpu: 1. This tells the scheduler nothing about the device, not its VRAM, compute capability, or whether it's a full 80 GB GPU or a 5 GB slice of a partitioned one. A training team reserving a full 80 GB GPU and an inference team using a 5 GB slice both consume "1" from the same quota. For platform teams trying to share a finite pool of accelerators across tenants with fundamentally different resource needs, this makes accurate accounting and fair sharing impossible.
Dynamic resource allocation (DRA), now generally available in Kubernetes 1.34, replaces this model entirely. Drivers advertise real device attributes—memory, compute capability, topology—through structured APIs, and workloads describe what they actually need rather than asking for a raw count. But DRA solves discovery and allocation. It doesn't address quota management, fair sharing, or preemption—the policies that determine which team's workload runs first when GPUs are scarce.
That's where Kueue comes in. Kueue's integration with DRA brings quota enforcement, borrowing, preemption, and admission fair sharing to DRA-managed devices. Red Hat build of Kueue 1.4 on Red Hat OpenShift makes enabling DRA support straightforward. The operator detects your cluster's DRA capabilities and enables the corresponding Kueue feature gates automatically. Administrators configure device class mappings in the Kueue custom resource, and the operator takes care of the rest. This article demonstrates how Kueue manages DRA device quota, from whole-device counting to memory-accurate partition accounting, and looks at where the integration is headed next.
DRA quota management in Kueue
Consider a platform team managing a shared GPU cluster for multiple data science teams. Without DRA integration, Kueue has no visibility into GPU requests made through ResourceClaimTemplates. Workloads get admitted without any GPU quota checks, and teams can consume more than their fair share of accelerators.
Device-level quota for whole GPUs
The foundation of Kueue's DRA integration solves this issue. In DRA, hardware drivers categorize devices into DeviceClasses. For example, NVIDIA's driver creates gpu.nvidia.com for whole GPUs. Workloads request devices by referencing a ResourceClaimTemplate that specifies a DeviceClass and a device count. Administrators configure a mapping between DeviceClass names and logical quota resource names in the Kueue custom resource:
spec:
config:
resources:
deviceClassMappings:
- name: nvidia-gpu
deviceClassNames:
- gpu.nvidia.comTo submit a workload requesting a GPU, users create a ResourceClaimTemplate referencing the DeviceClass, optionally with a CEL selector to target specific hardware:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: single-gpu
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.nvidia.com
count: 1
selectors:
- cel:
expression: "device.attributes['gpu.nvidia.com'].productName == 'A100'"The Kueue operator reads the ResourceClaimTemplate, resolves the DeviceClass to the nvidia-gpu quota resource, and charges the device count against ClusterQueue quota just like CPU or memory. All of Kueue's existing machinery (borrowing between queues, priority-based preemption, admission fair sharing) works with these DRA resources by default.
When a team submits a training job requesting GPUs, Kueue checks the ClusterQueue quota, admits if capacity is available, and suspends if not. When the job finishes, quota is freed and the next team's workload proceeds. A higher-priority job can preempt lower-priority ones to reclaim GPU quota. Teams can borrow GPU capacity from other queues within a cohort when it's idle.
This capability is production-ready. For step-by-step configuration, see the Red Hat build of Kueue DRA documentation.
Counter-based quota for partitionable GPUs
Consider a team managing a pool of GPUs shared between a training team and an inference team. The training team runs large jobs on whole GPUs. The inference team runs dozens of small models that might only need, say, 10 GB of GPU memory (for example, a 1g.10gb MIG profile on an NVIDIA GPU). Under device-count quota, both a whole 80 GB GPU and a 10 GB partition consume "1 device". The inference team's small workloads eat the same quota as the training team's full GPUs—making fair sharing impossible.
Partitionable devices solve this by shifting from device counts to counter-based quota. DRA drivers already expose how much memory each device (whole or partitioned) actually consumes. Kueue reads this counter data and charges GPU memory against ClusterQueue quota instead of device count. Administrators express quota in memory units (say, 800Gi total) and each workload is charged proportionally to the memory its partition actually consumes. The training team's whole-GPU job charges ~80 GB, and the inference team's partition charges ~10 GB. Both teams draw from the same quota pool, but the accounting reflects reality.
On OpenShift, Red Hat build of Kueue auto-detects the Kubernetes DRAPartitionableDevices feature gate and enables counter-based charging when sources are configured in the Kueue custom resource:
spec:
config:
resources:
deviceClassMappings:
- name: gpu-memory
deviceClassNames:
- gpu.nvidia.com
- mig.nvidia.com
sources:
- type: Counter
counter:
name: memory
driver: gpu.nvidia.comTo request a specific GPU partition, users create a ResourceClaimTemplate with a CEL selector targeting the desired profile, and reference it from their Job:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: gpu-partition
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: mig.nvidia.com
count: 1
selectors:
- cel:
expression: "device.attributes['gpu.nvidia.com'].profile == '1g.10gb'"Kueue reads the ResourceClaimTemplate, resolves the DeviceClass, reads the counter data for the matched partition, and charges the appropriate GPU memory against quota.
For step-by-step configuration, see the Red Hat build of Kueue DRA documentation.
What's next
DRA support in Kueue is moving fast. Support for partitionable devices has graduated to Beta in Kueue 0.19 and is now enabled by default, meaning counter-based GPU memory quota no longer requires explicitly enabling a feature gate. Expect this to land in a future version of Red Hat build of Kueue.
Extended resources is another approach to DRA quota management that simplifies adoption significantly. Instead of writing ResourceClaimTemplates, teams keep using the familiar resources.requests: {nvidia.com/gpu: 1} syntax they already use with device plugins. Kueue auto-discovers DRA-backed DeviceClasses and tracks quota automatically, with no deviceClassMappings configuration needed. When both extended resources and ResourceClaimTemplates are used in the same cluster, Kueue unifies quota across both paths, preventing double-counting. For teams planning their move from device plugins to DRA, this is one to watch.
Beyond what's shipping, the upstream Kueue community is exploring consumable capacity—the complement to partitionable devices. Where partitionable devices handle static GPU partitioning through MIG with slices fixed at configuration time, consumable capacity targets dynamic sharing through time-slicing and MPS, where multiple workloads run concurrently on the same physical GPU, each consuming a fraction of its compute and memory. Capacity-based quota would charge workloads based on the actual fraction they consume, preventing over-admission while improving utilization for inference and smaller training workloads.
Another area is closing the gap between quota and scheduling. Today, Kueue admits workloads based on aggregate quota: "Does this ClusterQueue have enough GPU memory?" But quota availability doesn't guarantee that pods can actually schedule on real nodes. A workload might hold quota while pods sit pending because no single node has the requested DRA devices available. The upstream community is working on integrating the scheduler-library, a new Kubernetes SIG-scheduling project that provides in-memory scheduling simulation using the actual kube-scheduler plug-ins. This library would let Kueue verify how feasible it would be to schedule a workload, checking not just quota but whether DRA devices are actually allocatable on specific nodes and whether gang placement is possible across all PodSets.
These are upstream developments and will continue to flow downstream through Red Hat build of Kueue as they mature.