Why Jev went viral, what "System One" models are good for, and how the vLLM community turned DiffusionGemma, already a Red Hat AI validated model, into an open, self-hostable decision engine.
A new shape of model
Over the last couple of weeks, a new kind of model took over developer feeds. Jev is the first model from TypeSafe AI, a San Francisco startup founded by former OpenAI researcher Diogo Almeida. Jev doesn't chat or write. It takes unstructured state in and returns typed, probabilistic decisions out, like a function call backed by frontier intelligence.
The API has 3 question types. A Choice question picks 1 option from a set, a Score question places input on an ordered scale, and a Noul question is a yes-or-no. The name comes from the Bernoulli distribution. Each answer comes back with a probability that your code can branch on.
TypeSafe calls this category "System One," after Daniel Kahneman's fast, intuitive System 1 thinking: The model picks immediately from defined options instead of deliberating token by token.
Why this resonated
Enterprises have been making these kinds of decisions with AI for a while. Search and e-commerce teams routinely run small generative models with structured output as zero-shot classifiers. As one analysis put it, composing AI into software is not new; Jev's contribution is a model and API built only for that role, with low latency, low cost, typed answers, and probabilities as the normal output.
So why did it go viral? Three reasons stand out:
- Guaranteed structure. The answer is always one of the options you defined, so there's no JSON to repair and no free text to parse.
- Probabilities by default. Confidence scores let you set thresholds, escalate uncertain cases to a person or a larger model, and act automatically on confident ones.
- Speed and cost. One structured pass is much cheaper than a full decode loop. TypeSafe reports latency of 70–500 ms.
The use cases are the unglamorous, high-volume decisions inside every application: ticket routing, content moderation, risk scoring, agent branching, and guardrails. A person opens a chatbot a few times a day, but software could make thousands of tiny decisions in the background.
There's a catch for many enterprises. Jev is a hosted API in early access, and TypeSafe hasn't published weights, a parameter count, or a self-hosting option as of this writing. Regulated industries, air-gapped environments, and sovereign or public sector deployments can't send every routing decision to a third-party endpoint. They need the pattern, not the dependency.
The open path: Decisions on DiffusionGemma in vLLM
The vLLM community found that 1 open model already contains most of what a decision engine needs.
DiffusionGemma 26B-A4B is Google's block-diffusion language model, built on Gemma 4's mixture of experts (MoE) backbone with 26B total and 4B active parameters. An autoregressive large language model (LLM) writes left to right, 1 token at a time. DiffusionGemma instead works on a fixed-length "canvas" of tokens and fills in every position in parallel through iterative denoising.
That parallelism is what makes parallel decisions possible. vLLM PR #57250 adds a structured-read mode, which works like this:
- Seed the canvas. The client prefills the canvas with the answer template, such as
urgent: @/category: @/severity: @. Only the answer slots are left as noise. - Run 1 denoising step. The request is capped at a single step and marked read-only, so vLLM skips the extra work that full generation would do.
- Read the probability distribution at each slot. Every answer slot returns calibrated logprobs from a single forward pass. The top choice is the decision, and the entropy of the distribution is the confidence.
- Reread only when uncertain. If a slot's entropy is above a threshold, the client samples a few more reads and measures agreement. Confident answers return after a single read.
Answers must be single tokens so the canvas layout stays fixed. That's easy to handle on the client: map moderation_spam to B, for example. The same mechanism covers yes-or-no, multiple-choice, and scored questions, and several questions can be asked in 1 request.
Early numbers are encouraging. On a single DGX Spark, the PR author measured 8.7 req/s at 0.12 s with 1 request at a time. With 32 concurrent requests they measured 54 req/s at 0.58 s; each request answered 3 questions, for roughly 162 decisions/s. Google has also published a Cloud Run deployment of this approach running on vLLM, reporting about 35–60 ms single-step latency and 100–123 req/s at batch 32; at three questions per request this translates to 300+ decisions/s.
Try it: The vLLM recipe
Decisions currently need a nightly build. The vLLM recipe walks through 3 steps.
1. Start vLLM with a diffusion canvas sized for decisions. A 64-token canvas is enough room for a template with several questions.
podman run -d --name dgemma --gpus '"device=0"' --ipc=host \
-p 8000:8000 -p 8011:8011 \
-v "$HOME/.cache/huggingface:/root/.cache/huggingface" \
vllm/vllm-openai:nightly-e9757321527ca1ecd514c07c1418dd2c53da3d19 \
google/diffusiongemma-26B-A4B-it \
--served-model-name dgemma \
--diffusion-config '{"canvas_length":64}' \
--max-logprobs 32 \2. Start the example decision server. It converts a schema into a seeded canvas, so clients never have to build one by hand.
until curl -fsS localhost:8000/health >/dev/null; do sleep 2; done
podman exec -d dgemma python \
/vllm-workspace/examples/features/structured_diffusion/structured_server.py \
--upstream http://127.0.0.1:8000 \
--tokenizer google/diffusiongemma-26B-A4B-it --canvas 643. Ask a question. The example server exposes a Jev-compatible /v1/systemone endpoint, so existing Jev client code can point at it.
until curl -fsS localhost:8011/health >/dev/null; do sleep 2; done
curl -sS localhost:8011/v1/systemone -H 'content-type: application/json' -d '{
"model": "jev-latest",
"state": {"ticket": "Everything is down, demo at noon."},
"questions": {
"urgent": {"type": "noul", "instructions": "Needs a reply within the hour?"}
}
}'Keep a few operational details in mind:
- The
/v1/systemoneendpoint is an example server, not a standard vLLM API. The engine-level machinery (seeded canvases, step caps, read-only requests) is in vLLM core. The request format is left to the community, which is still settling on one and will release it as an experimental endpoint. - For standard generation, the recipe keeps
--max-num-seqslow (4 at a 256-token canvas) because the diffusion state buffers scale with batch size, canvas length, and Gemma's 262K vocabulary. Decisions use a much smaller canvas, which is why the PR's 32-way concurrency numbers are possible. Size your deployment against the variant and canvas length you actually plan to use. - You can also try the quantized variants of DiffusionGemma for similar capability with a lower memory footprint.
Running decision models on Red Hat AI and Red Hat AI Inference
DiffusionGemma is already a Red Hat AI validated model for multimodal workloads (handling text, image, and video inputs). Red Hat has tested it for its existing use cases on the Red Hat AI platform and published optimized checkpoints in the RedHatAI Hugging Face collection, including FP8-dynamic and NVFP4 variants. The NVFP4 variant cuts the memory footprint to roughly a third of BF16.
Customers can evaluate decision mode today on Red Hat AI Inference preview builds with confidence that the model architecture is validated on Red Hat infrastructure. It will land in our supported version shortly after.
Structured decisions follow a clear rollout path across the platform:
- Prototyping today: Decisions are available in the upstream vLLM nightly, which teams can run on Red Hat AI Inference and Red Hat OpenShift AI through a custom serving runtime, which is unsupported, for prototyping against their own data.
- Try it now via the unsupported Red Hat AI Inference preview image.
- Next: When structured-read support ships in a stable vLLM release, Red Hat AI Inference Server will pick it up in stages, starting with a preview and hardening toward a supported endpoint based on customer feedback.
- Example server Developer Preview in 3.6 general availability (GA), assuming vLLM >= 0.31.0 is picked up.
- Hardened endpoint: timing and support level to be determined, based on feedback.
For customers who can't send data to a hosted API, such as regulated industries, air-gapped sites, and sovereign or public sector deployments, this is the key point. The model is already validated, the weights are open, and the decision engine runs on hardware you control.
The bigger picture
Decision models aren't a replacement for LLMs. They're a new building block that sits beside them. The pattern many teams will adopt is a fast decision model in the request path for routing, gating, and classification, with a generative model behind it for work that needs language and reasoning. Whatever interface the community settles on, the aim is for vLLM to run it in the open. As new models are developed, we will likely see many more decision-style models and one of our goals is to ensure these models run great on vLLM.
Start running decision models on infrastructure you control. Test the vLLM recipe on Red Hat OpenShift AI today, explore PR #57250, download the FP8-quantized DiffusionGemma checkpoints, or read the Red Hat AI Inference guide to plan your deployment roadmap.