Until recently, running distributed training on Kubernetes meant picking a framework-specific custom resource definition (CRD): PyTorchJob for PyTorch, TFJob for TensorFlow, MPIJob for Message Passing Interface (MPI). Each had different semantics, different failure behavior, and different ways of configuring the same fundamentals: how many nodes, how they find each other, and what happens when something breaks.
Red Hat OpenShift AI 3.4 changes this. It introduces Kubeflow Trainer v2, a unified training API that replaces framework-specific CRDs like PyTorchJob, TFJob, and MPIJob with a single TrainJob resource. Instead of writing a different resource for each machine learning (ML) framework and wiring up primary addresses, environment variables, and torchrun commands by hand, you reference a prebuilt ClusterTrainingRuntime and specify what matters: your training code, how many nodes, and how many GPUs per node.
Under the hood, Kubeflow Trainer v2 translates each TrainJob into a JobSet—a Kubernetes-native API for managing a group of Jobs as a single unit. The JobSet Operator manages these resources on OpenShift: it sets up stable headless services so workers can find and communicate with each other, provides automatic failure recovery from saved checkpoints, and controls startup sequencing to help meet dependencies.
This post walks through the full integration end to end—NVIDIA GPUs, real data, real failure scenarios—validated in 2 phases.
Phase 1: Manual validation
We confirmed all infrastructure components integrate correctly. We deployed both a TrainJob (via the Kubeflow Trainer v2 API) and a raw JobSet resource, each requesting NVIDIA GPUs and running a distributed PyTorch training script. Both approaches successfully allocated GPUs, established inter-pod networking via headless services, and computed correct NVIDIA Collective Communications Library (NCCL) all-reduce results across 2 nodes.
Phase 2: Real-world dataset training
In phase 2, we moved beyond synthetic tests to train an embedding-based neural network on a publicly available fraud detection dataset. We ran training across 1, 2, and 4 NVIDIA A10G GPUs to benchmark Distributed Data Parallel (DDP) speedup and validate fault tolerance, checkpoint recovery, and GPU telemetry.
- Dataset: Credit Card Transactions Fraud Detection Dataset (Sparkov-generated simulated transactions) (1,296,675 transactions, ~0.58% fraud rate, Creative Commons Zero (CC0)-licensed)
- Source: Fraud Dataset Benchmark (arXiv:2208.14417)
The full repository includes all manifests, training scripts, and a step-by-step Jupyter notebook you can run in a Red Hat OpenShift AI workbench.
The distributed training stack
You create a TrainJob that references the built-in torch-distributed ClusterTrainingRuntime. The JobSet Operator provisions the Kubernetes Job, a headless service for stable Domain Name System (DNS)-based pod discovery, and the training pods. Each pod lands on an NVIDIA GPU node and uses torchrun with NCCL to coordinate DDP training.
The stack relies on several operators working together:
- Kubeflow Trainer v2 automatically configures
torchrun, node coordination, and environment variables at runtime via prebuiltClusterTrainingRuntimes. - The JobSet Operator provisions headless services for stable network endpoints, manages configurable failure policies, and supports multi-template jobs with different pod specs per role.
- NVIDIA GPU Operator manages GPU drivers, monitoring, and device plug-ins required to expose physical nvidia.com/gpu resources to the cluster.
- Node Feature Discovery detects hardware capabilities on cluster nodes and labels them to help schedule workloads on the correct hardware.
- cert-manager provisions Transport Layer Security (TLS) certificates for webhook servers, required by the JobSet Operator.
- Red Hat build of Kueue acts as the resource gatekeeper, managing quotas, fair sharing, and all-or-nothing gang scheduling for distributed workloads (not covered here but part of the production stack).
These components, shown in Figure 1, are part of the distributed workloads feature in Red Hat OpenShift AI.

Prerequisites
This proof of concept (POC) ran on Red Hat OpenShift Container Platform 4.21 with NVIDIA GPU worker nodes and Red Hat OpenShift AI 3.4. For full installation instructions, see Installing OpenShift AI Self-Managed. The following operators were installed from OperatorHub:
- Node Feature Discovery (NFD): Labels GPU nodes so the NVIDIA operator can target them
- NVIDIA GPU Operator: Provides drivers, device plug-in, and Data Center GPU Manager (DCGM) for GPU scheduling
- cert-manager Operator: Provisions TLS certificates for webhook servers; required by JobSet Operator
- JobSet Operator: Required by Kubeflow Trainer v2
JobSet Operator configuration
Once the operator status shows Succeeded, create the custom resource (CR) for the JobSet Operator to deploy the controller pods:
apiVersion: operator.openshift.io/v1
kind: JobSetOperator
metadata:
name: cluster
spec:
logLevel: Normal
operatorLogLevel: Normal
managementState: ManagedVerify cert-manager, cert-manager-cainjector, and cert-manager-webhook pods are all Running.
DataScienceCluster configuration
In your DataScienceCluster CR, verify the following components are set (other components like dashboard, kserve, and workbenches can be configured based on your needs—see Installing OpenShift AI components):
trainer:
managementState: Managed # Kubeflow Trainer v2 (TrainJob API)
trainingoperator:
managementState: Removed # legacy v1, not used in this POC
kueue:
managementState: Removed # optional, not required for this POCVerify that a training runtime is available:
oc get clustertrainingruntimeExpected output includes torch-distributed.
The TrainJob API and raw JobSet offer different levels of control. The TrainJob API handles the distributed setup automatically. The raw JobSet path gives you full control but requires manual configuration. This post covers both.
Validating the operator chain
Two minimal distributed training flows confirmed the full operator stack—from the TrainJob API to the physical GPUs—integrates correctly on Red Hat OpenShift AI 3.4. Both tests use the same minimal PyTorch NCCL script, which initializes a process group, runs an all_reduce across 2 GPUs, and prints the result.
TrainJob API
The TrainJob requires only a runtime reference, a command, and the number of nodes. The torch-distributed ClusterTrainingRuntime handles torchrun, environment variables, and pod coordination automatically.
apiVersion: trainer.kubeflow.org/v1alpha1
kind: TrainJob
metadata:
name: pytorch-trainer-validation
namespace: <your-namespace>
spec:
runtimeRef:
name: torch-distributed
kind: ClusterTrainingRuntime
trainer:
command: ["torchrun", "/workspace/scripts/train.py"]
numNodes: 2
resourcesPerNode:
requests:
nvidia.com/gpu: 1You can find the full YAML in 01-trainjob-validation/trainjob.yaml.
Creating the TrainJob and ConfigMap triggered the full operator chain: Kubeflow Trainer v2 translated the request into a JobSet, and the JobSet Operator generated the Job, headless service, and 2 worker pods (Figure 2).

Raw JobSet
The second flow deployed a native JobSet directly, bypassing the TrainJob API. This required manually configuring the distributed environment, including MASTER_ADDR, MASTER_PORT, WORLD_SIZE, and RANK, and the full torchrun command.
apiVersion: jobset.x-k8s.io/v1alpha2
kind: JobSet
metadata:
name: pytorch-jobset-validation
namespace: <your-namespace>
spec:
replicatedJobs:
- name: workers
template:
spec:
parallelism: 2
completions: 2
backoffLimit: 0
template:
spec:
containers:
- name: pytorch
image: registry.redhat.io/rhoai/odh-training-cuda128-torch28-py312-rhel9:v3.0
command:
- /bin/sh
- -c
- |
torchrun \
--nproc_per_node=1 \
--nnodes=2 \
--rdzv_id=100 \
--rdzv_backend=c10d \
--rdzv_endpoint=$MASTER_ADDR:$MASTER_PORT \
/workspace/train.py
env:
- name: MASTER_ADDR
value: "pytorch-jobset-validation-workers-0-0.pytorch-jobset-validation"
- name: MASTER_PORT
value: "29500"
- name: WORLD_SIZE
value: "2"
- name: RANK
valueFrom:
fieldRef:
fieldPath: metadata.annotations['batch.kubernetes.io/job-completion-index']
resources:
requests:
nvidia.com/gpu: "1"You can find the full YAML in 02-jobset-validation/jobset.yaml.
Despite the manual setup, the result was identical (Figure 3)—same all_reduce = 1.0, same GPU allocation and NCCL communication.

The takeaway: Both paths successfully run distributed PyTorch training on NVIDIA GPUs. TrainJob simplifies orchestration with zero manual configuration, while raw JobSet resources provide full control with more manual setup.
Moving to a real workload: Fraud detection
With the infrastructure validated, the next step is a real, GPU-bound workload to benchmark multi-GPU speedup, shared storage, fault recovery, and GPU telemetry.
The dataset is the Credit Card Transactions Fraud Detection Dataset from the Fraud Dataset Benchmark collection (Sparkov-generated simulated transactions): 1,296,675 simulated credit card transactions with an approximate 0.58% fraud rate (CC0-licensed). The model is an embedding-based tabular neural network—learned embeddings for categorical fields, standardized numeric features, and a multilayer perceptron (MLP) classification head—trained with PyTorch DDP via torchrun across 1, 2, and 4 GPUs.
We deployed MinIO in-namespace as S3-compatible shared storage, sharing the dataset and checkpoints across multiple pods running the training in parallel. You can find the training script, data preparation code, and all TrainJob YAMLs in 03-fraud-detection/.
The training script checkpoints to MinIO in 2 ways:
- Periodic: Automatically every 200 steps during training. Used for recovery after a hard pod kill, where the pod has no warning and can't save anything before dying.
- On shutdown: When all pods receive a graceful stop signal (for example, suspend/resume), they save 1 final checkpoint before exiting, minimizing lost progress.
We validated both in the fault tolerance tests.
DDP speedup
We trained the same model for 10 epochs on 1, 2, and 4 NVIDIA A10G GPUs using the torch-distributed ClusterTrainingRuntime via TrainJob. Scaling from 1 to 4 GPUs requires changing the numNodes:
# 1-GPU: numNodes: 1
# 2-GPU: numNodes: 2
# 4-GPU: numNodes: 4
trainer:
command: ["/bin/sh", "-c"]
args:
- |
torchrun /workspace/scripts/train.py \
--data-dir=/mnt/local-cache/data \
--checkpoint-dir=/mnt/local-cache/checkpoints \
--checkpoint-prefix=checkpoints/4gpu \
--no-resume \
--epochs=10
numNodes: 4
resourcesPerNode:
requests:
nvidia.com/gpu: 1You can find the full YAML in trainjob-4gpu.yaml. The TrainJob defines what resources are needed.
The operator chain handles the rest:
- JobSet Operator creates a headless service so pods can discover each other by DNS name, and creates the pods on GPU nodes. This is the orchestration layer—it helps set up the distributed environment correctly before any training begins.
train.pyruns inside each pod and does the actual GPU work—downloads the dataset, connects to the other pods via the network configured byJobSet, splits the training data across GPUs, and trains the model. Each GPU processes its own batch simultaneously, and they synchronize after each step. After 10 epochs, rank 0 evaluates the model and reports the result.
Near-linear scaling—perfect scaling would be 2.00x and 4.00x, we got 1.94x and 3.80x, which is 95% efficiency at 4 GPUs (Figure 4). Training time dropped from more than 15 minutes on a single GPU to 4 minutes on 4. Model quality (Area Under the Receiver Operating Characteristic Curve (AUC-ROC)) remained stable across all splits, confirming DDP distributed the data correctly across GPUs.

Freeing GPUs without losing training progress
GPU training jobs can run for hours or days, tying up expensive accelerators that other teams or workloads need. TrainJob supports suspending a running job to pause training and free up cluster resources. If checkpointing is configured, the job saves the training state before terminating the pods. On resume, the job loads the latest checkpoint and continues from where it stopped.
Unlike the benchmark runs (which use --no-resume to start fresh), this scenario uses a separate TrainJob with checkpoint resume enabled. We suspended the running 2-GPU TrainJob mid-training to validate the JobSet Operator correctly coordinates stopping and restarting all GPU pods as a group:
oc patch trainjob fraud-detection-suspend-resume --type=merge -p '{"spec":{"suspend":true}}'The cluster tore down all GPU pods together. To resume:
oc patch trainjob fraud-detection-suspend-resume --type=merge -p '{"spec":{"suspend":false}}'The cluster re-created both pods together on resume. Training logs confirmed checkpoint recovery (Resumed from checkpoint-step-00000634.pt at epoch 1, step 634) and training completed normally.
GPU telemetry
During a 2-GPU training run, DCGM Exporter metrics (queried through OpenShift's Thanos-Querier) confirmed active GPU usage: compute utilization averaged 17% to 57% with peaks at 87% to 97%, and each GPU consumed around 440 MiB of memory. Before training, utilization was 0%—the workload is genuinely GPU-bound.
Orchestrating multi-stage pipelines with JobSet
The TrainJob API simplifies distributed training but abstracts away the underlying JobSet. This showcases features available in JobSet that go beyond what the TrainJob API covers—dependency ordering, granular failure policies, coordinator-based rendezvous, and automatic volume provisioning—using the same fraud detection workload.
In the previous test, data-prep and training ran as 2 separate commands. Here we combine them into a single raw JobSet (fraud-detection-full)—same scripts, same workload, but now the JobSet Operator manages the entire pipeline:
data-prepgroup (1 CPU pod): Downloads the dataset and uploads it to MinIOtraininggroup (2 GPU pods): Runstorchrun train.pywith DDP
You can find the full JobSet YAML in 04-jobset-features/fraud-detection-jobset.yaml.
Multi-template jobs
JobSet models a distributed workload as a group of Jobs, allowing different pod templates for distinct groups of pods—something a single Kubernetes Job can't do. The 2 groups shown in Figure 5 have completely different pod specs: data-prep is CPU-only, training has GPU requests and tolerations. One oc apply creates the entire pipeline.

dependsOn and successPolicy
Without dependency ordering, training pods would start immediately and fail because the dataset isn't ready—wasting GPU allocation time. Without a success policy, the JobSet could be marked complete when data-prep finishes, before training even runs. The training group in the JobSet declares 2 policies:
dependsOn:
- name: data-prep
status: Complete
successPolicy:
operator: All
targetReplicatedJobs:
- trainingdependsOn declares training can't start until data-prep succeeds. successPolicy declares the JobSet is complete only when the training group finishes (not data-prep alone). The watch loop output shows both working together:
Check 2: data-prep Running - training not created yet (dependsOn)
Check 5: data-prep Succeeded - training Running (pods appeared)
Check 6: JobSet complete: False - training still running (successPolicy)
Check 14: JobSet complete: True - both groups SucceededPipeline ordering and completion criteria work together—training waits for data, and the JobSet waits for training.
Coordinator
Distributed training requires all nodes to discover a single rendezvous endpoint to coordinate. Without the coordinator field, this requires hardcoding the endpoint or adding custom init containers to resolve it at runtime.
coordinator:
replicatedJob: training
jobIndex: 0
podIndex: 0The coordinator field tells JobSet to label all pods with jobset.sigs.k8s.io/coordinator, pointing to the rank-0 training pod's DNS name. This lets any pod discover the rendezvous endpoint automatically. When using the TrainJob API, this is configured automatically by the runtime.
In this deployment, all training pods received the coordinator label: fraud-detection-full-training-0-0.fraud-detection-full—the rank-0 training pod's DNS name used as the torchrun rendezvous endpoint.
Failure policies
Data preparation and training can fail for different reasons and should be handled accordingly. A data preparation failure means the dataset is unavailable, so there is no reason to start training. A training failure—such as a killed pod, out of memory (OOM), or node preemption—is often transient and worth retrying. JobSet 's failure policy allows different actions to be configured for each stage.
Two failure rules handle different failure scenarios:
failurePolicy:
maxRestarts: 3
rules:
- name: fail_on_data_prep_error
action: FailJobSet
targetReplicatedJobs:
- data-prep
- name: restart_on_training_failure
action: RestartJobSet
targetReplicatedJobs:
- trainingTo test FailJobSet, we deployed a separate JobSet where data-prep deliberately fails with exit 1. The result:
NAME TERMINALSTATE RESTARTS
fraud-fail-dp Failed 0 ← JobSet failed immediately
NAME STATUS
fraud-fail-dp-data-prep-0-0-j7xv6 Error ← only data-prep ran, no training pods createdThe JobSet immediately failed and training never started. No wasted GPU time on missing data.
For RestartJobSet, we killed a training pod mid-run with oc delete pod --force. The cluster re-created all training pods, and the restarted run completed faster by resuming from the last MinIO checkpoint instead of starting from scratch.
volumeClaimPolicies
The volumeClaimPolicies feature auto-provisioned a 5 GiB PersistentVolumeClaim (PVC) for data-prep scratch space on oc apply, and auto-deleted it on JobSet deletion. On a cluster with ReadWriteMany (RWX) storage, this feature would let all training pods mount the same auto-provisioned PVC directly. On our ReadWriteOnce (RWO)-only cluster, the auto-provisioned PVC could only serve single-pod use, so we handled multi-node data sharing through MinIO instead.
Wrap up
In this post, we walked through the full distributed training stack on Red Hat OpenShift AI 3.4—from basic operator validation to multi-GPU benchmarking and fault recovery on a real dataset. Here's what stood out:
- Simplicity: The
TrainJobAPI handlestorchrun, environment variables, and pod coordination automatically. Define your training code, node count, and GPU requirements—no manualMASTER_ADDR, no entrypoint scripts, no framework-specific CRDs. - Flexibility: When your workload needs more control, raw JobSets give you multi-stage pipelines, startup sequencing, and differentiated failure handling—all on the same underlying operator stack.
- Reliability: Both paths delivered near-linear DDP scaling at 95% efficiency, successful checkpoint recovery after suspend/resume and hard pod kills, and GPU utilization confirmed via DCGM Exporter during training runs.
In short, Red Hat OpenShift AI 3.4 provides a production-ready platform for distributed GPU training. Whether you use the high-level TrainJob API or the low-level JobSet path, the operator chain handles the orchestration so you can focus on the model.
Next steps
- Try it yourself: Run the POC notebook in a Red Hat OpenShift AI workbench to deploy, verify, and clean up each test case end to end. The repository includes all manifests and training scripts ready to apply. Full setup details are in the README.
- Adapt it to your workload: The fraud detection model is 1 example. The same
TrainJob+JobSetpatterns work for any PyTorch DDP workload—swap the training script, adjust the GPU count. - Get started: Try Red Hat OpenShift AI.
Learn more
If you want to explore distributed training on Red Hat OpenShift AI, learn more about the tech stack: