Skip to main content
Redhat Developers  Logo
  • AI

    Get started with AI

    • Red Hat AI
      Accelerate the development and deployment of enterprise AI solutions.
    • AI learning hub
      Explore learning materials and tools, organized by task.
    • AI interactive demos
      Click through scenarios with Red Hat AI, including training LLMs and more.
    • AI/ML learning paths
      Expand your OpenShift AI knowledge using these learning resources.
    • AI quickstarts
      Focused AI use cases designed for fast deployment on Red Hat AI platforms.
    • No-cost AI training
      Foundational Red Hat AI training.

    Featured resources

    • OpenShift AI learning
    • Open source AI for developers
    • AI product application development
    • Open source-powered AI/ML for hybrid cloud
    • AI and Node.js cheat sheet

    Red Hat AI Factory with NVIDIA

    • Red Hat AI Factory with NVIDIA is a co-engineered, enterprise-grade AI solution for building, deploying, and managing AI at scale across hybrid cloud environments.
    • Explore the solution
  • Learn

    Self-guided

    • Documentation
      Find answers, get step-by-step guidance, and learn how to use Red Hat products.
    • Learning paths
      Explore curated walkthroughs for common development tasks.
    • Guided learning
      Receive custom learning paths powered by our AI assistant.
    • See all learning

    Hands-on

    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.
    • Interactive labs
      Learn by doing in these hands-on, browser-based experiences.
    • Interactive demos
      Click through product features in these guided tours.

    Browse by topic

    • AI/ML
    • Automation
    • Java
    • Kubernetes
    • Linux
    • See all topics

    Training & certifications

    • Courses and exams
    • Certifications
    • Skills assessments
    • Red Hat Academy
    • Learning subscription
    • Explore training
  • Build

    Get started

    • Red Hat build of Podman Desktop
      A downloadable, local development hub to experiment with our products and builds.
    • Developer Sandbox
      Spin up Red Hat's products and technologies without setup or configuration.

    Download products

    • Access product downloads to start building and testing right away.
    • Red Hat Enterprise Linux
    • Red Hat AI
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Featured

    • Red Hat build of OpenJDK
    • Red Hat JBoss Enterprise Application Platform
    • Red Hat OpenShift Dev Spaces
    • Red Hat Developer Toolset

    References

    • E-books
    • Documentation
    • Cheat sheets
    • Architecture center
  • Community

    Get involved

    • Events
    • Live AI events
    • Red Hat Summit
    • Red Hat Accelerators
    • Community discussions

    Follow along

    • Articles & blogs
    • Developer newsletter
    • Videos
    • Github

    Get help

    • Customer service
    • Customer support
    • Regional contacts
    • Find a partner

    Join the Red Hat Developer program

    • Download Red Hat products and project builds, access support documentation, learning content, and more.
    • Explore the benefits

Run decision models on vLLM and Red Hat AI using DiffusionGemma

Decision models like Jev are having a moment. Here's how to run one on vLLM and Red Hat AI today.

September 28, 2026
Lucas Wilkinson Rob Greenberg
Related topics:
Artificial intelligence
Related products:
Red Hat AI

    Why Jev went viral, what "System One" models are good for, and how the vLLM community turned DiffusionGemma, already a Red Hat AI validated model, into an open, self-hostable decision engine.

    A new shape of model

    Over the last couple of weeks, a new kind of model took over developer feeds. Jev is the first model from TypeSafe AI, a San Francisco startup founded by former OpenAI researcher Diogo Almeida. Jev doesn't chat or write. It takes unstructured state in and returns typed, probabilistic decisions out, like a function call backed by frontier intelligence.

    The API has 3 question types. A Choice question picks 1 option from a set, a Score question places input on an ordered scale, and a Noul question is a yes-or-no. The name comes from the Bernoulli distribution. Each answer comes back with a probability that your code can branch on.

    TypeSafe calls this category "System One," after Daniel Kahneman's fast, intuitive System 1 thinking: The model picks immediately from defined options instead of deliberating token by token.

    Why this resonated

    Enterprises have been making these kinds of decisions with AI for a while. Search and e-commerce teams routinely run small generative models with structured output as zero-shot classifiers. As one analysis put it, composing AI into software is not new; Jev's contribution is a model and API built only for that role, with low latency, low cost, typed answers, and probabilities as the normal output.

    So why did it go viral? Three reasons stand out:

    • Guaranteed structure. The answer is always one of the options you defined, so there's no JSON to repair and no free text to parse.
    • Probabilities by default. Confidence scores let you set thresholds, escalate uncertain cases to a person or a larger model, and act automatically on confident ones.
    • Speed and cost. One structured pass is much cheaper than a full decode loop. TypeSafe reports latency of 70–500 ms.

    The use cases are the unglamorous, high-volume decisions inside every application: ticket routing, content moderation, risk scoring, agent branching, and guardrails. A person opens a chatbot a few times a day, but software could make thousands of tiny decisions in the background.

    There's a catch for many enterprises. Jev is a hosted API in early access, and TypeSafe hasn't published weights, a parameter count, or a self-hosting option as of this writing. Regulated industries, air-gapped environments, and sovereign or public sector deployments can't send every routing decision to a third-party endpoint. They need the pattern, not the dependency.

    The open path: Decisions on DiffusionGemma in vLLM

    The vLLM community found that 1 open model already contains most of what a decision engine needs.

    DiffusionGemma 26B-A4B is Google's block-diffusion language model, built on Gemma 4's mixture of experts (MoE) backbone with 26B total and 4B active parameters. An autoregressive large language model (LLM) writes left to right, 1 token at a time. DiffusionGemma instead works on a fixed-length "canvas" of tokens and fills in every position in parallel through iterative denoising.

    That parallelism is what makes parallel decisions possible. vLLM PR #57250 adds a structured-read mode, which works like this:

    1. Seed the canvas. The client prefills the canvas with the answer template, such as urgent: @ / category: @ / severity: @. Only the answer slots are left as noise.
    2. Run 1 denoising step. The request is capped at a single step and marked read-only, so vLLM skips the extra work that full generation would do.
    3. Read the probability distribution at each slot. Every answer slot returns calibrated logprobs from a single forward pass. The top choice is the decision, and the entropy of the distribution is the confidence.
    4. Reread only when uncertain. If a slot's entropy is above a threshold, the client samples a few more reads and measures agreement. Confident answers return after a single read.

    Answers must be single tokens so the canvas layout stays fixed. That's easy to handle on the client: map moderation_spam to B, for example. The same mechanism covers yes-or-no, multiple-choice, and scored questions, and several questions can be asked in 1 request.

    Early numbers are encouraging. On a single DGX Spark, the PR author measured 8.7 req/s at 0.12 s with 1 request at a time. With 32 concurrent requests they measured 54 req/s at 0.58 s; each request answered 3 questions, for roughly 162 decisions/s. Google has also published a Cloud Run deployment of this approach running on vLLM, reporting about 35–60 ms single-step latency and 100–123 req/s at batch 32; at three questions per request this translates to 300+ decisions/s.

    Try it: The vLLM recipe

    Decisions currently need a nightly build. The vLLM recipe walks through 3 steps.

    1. Start vLLM with a diffusion canvas sized for decisions. A 64-token canvas is enough room for a template with several questions.

    podman run -d --name dgemma --gpus '"device=0"' --ipc=host \
      -p 8000:8000 -p 8011:8011 \
      -v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
      vllm/vllm-openai:nightly-e9757321527ca1ecd514c07c1418dd2c53da3d19 \
      google/diffusiongemma-26B-A4B-it \
      --served-model-name dgemma \
      --diffusion-config '{"canvas_length":64}' \
      --max-logprobs 32 \

    2. Start the example decision server. It converts a schema into a seeded canvas, so clients never have to build one by hand.

    until curl -fsS localhost:8000/health >/dev/null; do sleep 2; done
    
    podman exec -d dgemma python \
      /vllm-workspace/examples/features/structured_diffusion/structured_server.py \
      --upstream http://127.0.0.1:8000 \
      --tokenizer google/diffusiongemma-26B-A4B-it --canvas 64

    3. Ask a question. The example server exposes a Jev-compatible /v1/systemone endpoint, so existing Jev client code can point at it.

    until curl -fsS localhost:8011/health >/dev/null; do sleep 2; done
    
    curl -sS localhost:8011/v1/systemone -H 'content-type: application/json' -d '{
      "model": "jev-latest",
      "state": {"ticket": "Everything is down, demo at noon."},
      "questions": {
        "urgent": {"type": "noul", "instructions": "Needs a reply within the hour?"}
      }
    }'

    Keep a few operational details in mind:

    • The /v1/systemone endpoint is an example server, not a standard vLLM API. The engine-level machinery (seeded canvases, step caps, read-only requests) is in vLLM core. The request format is left to the community, which is still settling on one and will release it as an experimental endpoint.
    • For standard generation, the recipe keeps --max-num-seqs low (4 at a 256-token canvas) because the diffusion state buffers scale with batch size, canvas length, and Gemma's 262K vocabulary. Decisions use a much smaller canvas, which is why the PR's 32-way concurrency numbers are possible. Size your deployment against the variant and canvas length you actually plan to use.
    • You can also try the quantized variants of DiffusionGemma for similar capability with a lower memory footprint.

    Running decision models on Red Hat AI and Red Hat AI Inference

    DiffusionGemma is already a Red Hat AI validated model for multimodal workloads (handling text, image, and video inputs). Red Hat has tested it for its existing use cases on the Red Hat AI platform and published optimized checkpoints in the RedHatAI Hugging Face collection, including FP8-dynamic and NVFP4 variants. The NVFP4 variant cuts the memory footprint to roughly a third of BF16.

    Customers can evaluate decision mode today on Red Hat AI Inference preview builds with confidence that the model architecture is validated on Red Hat infrastructure. It will land in our supported version shortly after.

    Structured decisions follow a clear rollout path across the platform:

    • Prototyping today: Decisions are available in the upstream vLLM nightly, which teams can run on Red Hat AI Inference and Red Hat OpenShift AI through a custom serving runtime, which is unsupported, for prototyping against their own data.
    • Try it now via the unsupported Red Hat AI Inference preview image.
    • Next: When structured-read support ships in a stable vLLM release, Red Hat AI Inference Server will pick it up in stages, starting with a preview and hardening toward a supported endpoint based on customer feedback.
      • Example server Developer Preview in 3.6 general availability (GA), assuming vLLM >= 0.31.0 is picked up.
      • Hardened endpoint: timing and support level to be determined, based on feedback.

    For customers who can't send data to a hosted API, such as regulated industries, air-gapped sites, and sovereign or public sector deployments, this is the key point. The model is already validated, the weights are open, and the decision engine runs on hardware you control.

    The bigger picture

    Decision models aren't a replacement for LLMs. They're a new building block that sits beside them. The pattern many teams will adopt is a fast decision model in the request path for routing, gating, and classification, with a generative model behind it for work that needs language and reasoning. Whatever interface the community settles on, the aim is for vLLM to run it in the open. As new models are developed, we will likely see many more decision-style models and one of our goals is to ensure these models run great on vLLM.

    Start running decision models on infrastructure you control. Test the vLLM recipe on Red Hat OpenShift AI today, explore PR #57250, download the FP8-quantized DiffusionGemma checkpoints, or read the Red Hat AI Inference guide to plan your deployment roadmap.

    Recent Posts

    • Upgrading to Red Hat JBoss Web Server 7: Key changes & impacts

    • How to rank fraud detection models using custom cost metrics

    • Run decision models on vLLM and Red Hat AI using DiffusionGemma

    • Managing edge solutions with Red Hat: A layered approach

    • Build golden path CI/CD workflows in Red Hat Developer Hub

    What’s up next?

    Learning Path Get started with vLLM feature share

    Get started with vLLM

    Learn how to compress, serve, and benchmark LLMs with vLLM.
    Red Hat Developers logo LinkedIn YouTube Twitter Facebook

    Platforms

    • Red Hat AI
    • Red Hat Enterprise Linux
    • Red Hat OpenShift
    • Red Hat Ansible Automation Platform
    • See all products

    Build

    • Developer Sandbox
    • Developer tools
    • Interactive tutorials
    • API catalog

    Quicklinks

    • Learning resources
    • E-books
    • Cheat sheets
    • Blog
    • Events
    • Newsletter

    Communicate

    • About us
    • Contact sales
    • Find a partner
    • Report a website issue
    • Site status dashboard
    • Report a security problem

    RED HAT DEVELOPER

    Build here. Go anywhere.

    We serve the builders. The problem solvers who create careers with code.

    Join us if you’re a developer, software engineer, web designer, front-end designer, UX designer, computer scientist, architect, tester, product manager, project manager or team lead.

    Sign me up

    Red Hat legal and privacy links

    • About Red Hat
    • Jobs
    • Events
    • Locations
    • Contact Red Hat
    • Red Hat Blog
    • Inclusion at Red Hat
    • Cool Stuff Store
    • Red Hat Summit
    © 2026 Red Hat

    Red Hat legal and privacy links

    • Privacy statement
    • Terms of use
    • All policies and guidelines
    • Digital accessibility
    Ask AI