Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,569 words · 5 segments analyzed
Sep 24, 2026 Typed decisions with a probability for every option, in a single forward pass: matching Jev's accuracy and speed with an LLM.Johannes HötterVP GrowthMarko Rosenmüller, PhDTechnical Lead AITL;DR: In this post, we show how an off-the-shelf LLM can make typed decisions in a single forward pass. This approach makes it possible to turn an LLM into a Jev-like decision model. We evaluate the approach using GLM-5.3-Flash running on Privatemode. Using a benchmark constructed from public data sets, we show that this setup delivers results that are on par with TypeSafe's Jev in terms of decision accuracy/correctness and speed. As a bonus, the setup with GLM-5.3-Flash enables typed decisions on images, which is not possible with Jev. How it worksPlaygroundBenchmark Why typed decisions Much of what software asks an LLM is a decision. "Which team should handle this ticket?", or "Does this contract clause belong in the liability section?". In such cases, software typically requires that the LLM's response follows a certain format like JSON and that it comes from a pre-defined set like "yes" and "no". Given the right instructions, LLMs can typically already fulfill this reliably. However, in the basic approach, speed and costs become an issue: For each decision, the LLM needs to write a whole JSON object, and a reasoning model may think for hundreds of tokens before that. Further, you also don't learn the confidence of the model (unless you explicitly ask it). All these aspects can matter a lot in practice and have so far prevented people from employing LLMs for decision making in high-volume/high-throughput scenarios. Specialized decision models (or "System One" models) like Jev and Laya are designed to address this. You pass in a piece of state and a set of named options, and you get back the chosen option together with a confidence value (i.e., probability) for each one. Turning an LLM into a decision model Initially, we asked ourselves if an LLM could be turned into a decision model with Jev-like properties. The short answer is: "yes". In the following, we show how it works. To understand our approach, it's important to understand how LLMs work: An LLM never writes text directly. Given a prompt, an LLM outputs a probability distribution over its entire vocabulary of tokens. In text generation, in the simplest case, the token with the highest probability is selected as the next token. The selected token then is appended to the prompt and the whole process repeats. As described above, this is costly and slow if you just want to set a few fields in a JSON object. Our core insight is that it's unnecessary to have the LLM predict the whole JSON object, as we already know its shape. We're only interested in the LLM's typed judgement for a given input. We realized that it's possible to craft prompts so that we get the typed judgement in a single run of the LLM — with no fine-tuning, on the model exactly as it ships. This is the difference to Jev and Laya, which are models trained for the purpose. The basic steps are as follows: Number the options. The state, the question, and the output options go into the prompt as JSON, with an index on every option. The instruction asks the model to answer with choice_index: followed by an index. Prefill the answer. The prompt ends with choice_index:. Consequently, the first token the model produces will be an index into the pre-defined options. Evaluate the output. Rather than reading the token the model emits, we read the probabilities it assigned to all option indexes at that single position. Normalized over the options, these give a probability for each answer, and we simply pick the most probable one. End the prompt inside the answeruser{"state": "I was charged twice for my order.", "question": "Which team?", "options": [{"index": 0, "name": "payments"}{"index": 1, "name": "complaints"}{"index": 2, "name": "technical"}]}assistantchoice_index:Read one row of logits0 1 2…whole vocabularymax_tokens: 1temperature: 0allowed_token_idsKeep the options, renormalize0 payments62.1%1 complaints37.7%2 technical0.2%Answer paymentsconfidence 0.39, where 1 means certain and 0 means the options are equally likelyThe prompt numbers the options and it ends with the assistant’s answer already begun as choice_index:, so the next token is the index. A mask on the vocabulary allows only the option indexes, the API returns their log probabilities, and normalizing them over the options gives a probability for each answer. The logit values in the middle panel are illustrative. We implemented the above steps for GLM-5.3-Flash running on vLLM (in Privatemode). We use the /chat/completions endpoint with continue_final_message and add_generation_prompt: false, because these let the model continue the prefilled assistant turn from step 2 instead of starting a new one. They also let us pass images next to the text, which is what makes typed decisions on images possible. For vLLM and GLM-5.3-Flash, we found the following details to matter: vLLM's allowed_token_ids can be used to limit the LLM's output vocabulary only to allowed options. It drops every other token to -inf. We set it, but it is a guardrail rather than a requirement. top_logprobs is not enough for step 3. It reports the distribution before the restriction is applied, so formatting tokens such as a leading space take up the top slots, and some options drop off the list and appear to have a probability of zero. vLLM's logprob_token_ids solves this: it returns the log probability of exactly the token ids you ask for. The token ids of the indexes depend on the model's tokenizer. Digits aren't always single tokens. GLM-5.3-Flash, for example, has a single token for 12. Rather than shipping a model-specific tokenizer, the library gets the token ids from the server, which keeps it simple to use with any model: sending a prompt to /completions with echo returns its exact tokenization by the model that is actually serving. You can find our implementation in the below repository. edgelesssys/privatemode-decisionsThe Python library: token oracle, prompt, masking and renormalization, against any vLLM-backed endpoint. Try it The playground below runs GLM-5.3-Flash queried with the above setup on Privatemode, directly from your browser. Pick one of the examples, among them a scanned invoice and a question that depends on your local time, or write your own questions and add images. Each answer comes back as a distribution over its options, typically within a few hundred milliseconds. The distribution is often very useful, e. g., to decide whether to include a human-in-the-loop. The model solves most classic trick questions, but not all of them. Benchmark results We evaluated our approach using a custom benchmark, which is available in the below repository. edgelesssys/privatemode-decisions-benchmarkThe benchmark: methodology, frozen dataset specs, harness and aggregation. Every number in this post can be recomputed from it. We compared three systems on 29 public, labeled datasets: GLM-5.3-Flash hosted on Privatemode and queried with the technique above, TypeSafe's Jev, and Convai's Laya.
The datasets have between 2 and 151 options and cover intent routing, sentiment, topic classification, moderation, entailment, question answering, legal text, and scanned documents. Both English and German text is included in the corpus. All three systems receive the same state, the same option names in the same order, and the same instruction. We ran each dataset twice. Even at temperature 0, our GLM-5.3-Flash and Jev changed up to 3.5% of their answers between identical runs: temperature 0 removes the randomness from sampling, but batching and floating-point arithmetic still keep a forward pass on a busy server from being bit-reproducible. We therefore treat smaller differences as noise.
We ran Jev and Laya with their default settings and did not tune our prompt on these datasets. Accuracy Compared across the 28 text datasets, GLM-5.3-Flash and Jev are on par. Each is more accurate on 10 datasets; on the remaining 8, the two are within one percentage point of each other. The median gap is 0.7 percentage points in Jev's favor, which is not statistically significant (p = 0.64).
Laya, a model with 421 million parameters that we ran locally, scores lower than both on most datasets. Its median gap is 13 to 15 percentage points, which is statistically significant (p < 0.001). The number of options has a larger effect on accuracy than the choice between Jev and GLM-5.3-Flash.
Which one is more accurate, dataset by dataset?10 GLM-5.3-Flash more accurate8 about the same10 Jev more accurateThe typical gap is 0.7 percentage points, slightly in Jev’s favor. Two equally accurate systems would show a gap at least this large in about 6 out of 10 comparisons, so it is well within chance.Accuracy by number of optionsGLM-5.3-FlashJevLayawith reasoningembedding similarity20%40%60%80%100%26 datasets3–68 datasets7–208 datasets21–804 datasets81+1 datasetHover or tap a mark for its numbers. Top: each square is one of the 28 datasets both hosted systems answer, colored by which of the two was more accurate on it; within one percentage point counts as the same, the variation between two identical runs. The chance estimate is a two-sided Wilcoxon signed-rank test over the per-dataset differences (p = 0.64). Bottom: mean accuracy by number of options, averaged only over datasets all three systems answered, so each point covers the same questions; the bands along the bottom are not evenly sized. The dashed and dotted lines are controls on the same datasets, both zero-shot like the rest: GLM-5.3-Flash allowed to reason before it answers, and plain embedding similarity with no decision model. Across datasets, the number of options changes along with everything else about the task.