Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 1,663 words · 6 segments analyzed
October 1, 2026Say you send an agent's next step to a decision model, and you only let it act on its own when it's at least 99% sure. Everything else goes to a person. That rule is only as good as the 99%. We gave GPT-6 Luna, the model behind OpenAI's new Decisions API, 3,600 reasoning problems and asked for exactly that kind of answer. When Luna said it was 99% sure or more, it was right 68% of the time. OpenAI announced the Decisions API at DevDay on September 29, in limited preview. It picks one answer from a list you define, so software can branch on it: classify a message, route a request, choose what an agent does next. The Decoder and Axios report it runs on "a version of GPT-6 Luna," OpenAI's low-cost model, and OpenAI's own chart puts it at 150 milliseconds a decision. It's OpenAI's answer to Jev. We'd already run Luna that way. In Hard-Decisions, our benchmark of decision models on multi-step logic, Luna answered the same problems as Jev, Kev and Laya, one request each, from a fixed list of options. That makes it a preview of a Luna-based decision model. It isn't the Decisions API itself: OpenAI uses a specialized version of Luna, and there's no documentation yet to test against. Why we care We've been routing on a model's confidence since 2023. For Call Criteria, LLMs score calls against hundreds of client scorecards, and people review the calls the system is unsure about. Everything depends on knowing which calls those are. When we read that confidence from the token log-probabilities of OpenAI's models, the numbers were so concentrated that they looked certain about nearly every answer. OpenAI's models have had confidence problems for as long as we've used them, and not the shy kind. Newer reasoning models often don't expose log-probabilities at all. So we built the confidence ourselves: extract the probability behind each label, check it against labeled answers, calibrate it, and only then route on it. Classification with Confidence shows that pipeline on GPT-4o-mini. That's the pain the Decisions API could end. What we're hoping for most is that OpenAI now treats calibration as a product: a confidence you can set a threshold on, documented and tested, instead of something withheld from the API, which we suspect, without proof, is partly about making the models harder to distill. A decision endpoint is where that would have to happen. So we wanted to know what the model underneath does today. Accuracy by the number of inference steps the answer needs, both ProofWriter tasks pooled. The test The problems come from ProofWriter, a dataset from the Allen Institute for AI. Each one is a short list of facts and if-then rules in plain English, and a statement to judge. Some answers can be read straight off the page. Others need five rules chained together. The dataset labels every problem with that number, its proof depth, which makes difficulty something you can measure instead of argue about. There are two versions of the task. In the open-world one the answer is true, false or unknown: unknown when the rules settle neither the statement nor its opposite. In the closed-world one, anything you can't prove is false. We sampled 1,800 problems for each, about 300 at every depth, and recomputed every gold answer with an independent solver. Luna got the same treatment as everyone else. One request per problem. The problem, the question and the meaning of each option, word for word what Jev receives. A strict JSON schema that only allows the three options. Reasoning off, which is the setting a fast decision API implies. We wrote our predictions down before any model answered. Accuracy: fine on the page, coin flips five steps in On problems you can answer by reading, Luna is nearly as good as Jev: 93% against 98% on the open-world task, and 98% against 99% on the closed one. Then it loses ground with every step. By five chained inferences Luna gets 46% and 45% right. On the closed-world task, that's a coin flip. Jev gets 81% and 89% at the same depth, on the same problems. Across all depths, Jev scored 83.8% and 89.3%; Luna, 64.1% and 65.0%. The gap at depth 5 is 35 and 45 points, with 95% intervals from 27 to 43 and from 38 to 51. That isn't a close call. You might be thinking that reasoning off is what sank Luna. Turning it on might help, but it would also make every decision slower and costlier than the one-shot job a decision API exists to do, and every model here got the same single shot. The open models we tested are stuck at depth 5 too: every model except Jev was at coin-flip accuracy by depth 5. Wrong when the answer is on the page Most of Luna's misses come from long chains, but not all of them. At depth 0, where the statement or its opposite is written in the text, Luna got 20 of 300 open-world problems wrong and 6 of 300 closed-world ones. At depth 1, one rule away, it missed 67 of 302 and 70 of 300. Here are the shortest of those misses at each depth, in each version of the task, with the confidence Luna gave when we asked again for log-probabilities and it gave the same answer: Depth 0, open world. The text says "The lion is not big." Is "The lion is big" true, false or unknown? It's false; the text says so.
Luna said unknown, 100% sure. Depth 0, closed world. The text says "Charlie is nice." Is "Charlie is not nice" true or false? False. Luna said true, 100% sure. Depth 1, open world. "The tiger is big. All big people are cold." Is "The tiger is cold" true? Yes, one rule away. Luna said unknown, 100% sure. Depth 1, closed world. "The rabbit is big. If someone is big then they chase the rabbit." Is "The rabbit does not chase the rabbit" true or false? False: the rabbit is big, so it chases the rabbit. Luna said true, 100% sure.
They're the clearest of Luna's misses; at depth 0 it's right 93% and 98% of the time. But a decision model that can be 100% sure Charlie is not nice, when the text says Charlie is nice, needs more than a threshold on its confidence. Can Luna tell you how sure it is? A decision model earns its keep when you can trust its confidence. Jev returns a probability for every option. Luna, through the regular API, returns text. But with reasoning off, OpenAI's API will also return log-probabilities: for each token Luna writes, the log of the probability it gave that token, plus the most likely alternatives. With reasoning on, it refuses: "'logprobs' is not supported with this model," which matches OpenAI's guide. The strict schema makes the reading clean. It only lets Luna reply {"answer":"<option>"}, so the token that spells the option carries the whole decision, with no "No" versus "no" to reconcile. Turn its log-probability into a probability and you have Luna's confidence in the answer it gave. Here's one real exchange.
The request, minus the long prompt: { "model": "gpt-6-luna", "reasoning_effort": "none", "max_completion_tokens": 1024, "response_format": { "type": "json_schema", "json_schema": { "name": "decision", "strict": true, "schema": { "type": "object", "properties": { "answer": { "type": "string", "enum": ["true", "false", "unknown"] } }, "required": ["answer"], "additionalProperties": false } } }, "logprobs": true, "top_logprobs": 5 } The problem, in its paraphrased form: Alan is young, round, and kind, but that doesn't mean he isn't also rough and cold at times, as well. … Young round people who are green are usually blue. … Kind people with rough skin are usually red because it's wind burn. If someone shows that they are red, then they are also showing that they are green. … Statement: Alan is not blue. And Luna's answer, with its log-probabilities: "content": "{\"answer\":\"unknown\"}", "logprobs": { "content": [ { "token": "{\"", "logprob": 0.0, "top_logprobs": [ { "token": "{\"", "logprob": 0.0 } ] }, { "token": "answer", "logprob": 0.0, "top_logprobs": [ { "token": "answer", "logprob": 0.0 } ] }, { "token": "\":\"", "logprob": 0.0, "top_logprobs": [ { "token": "\":\"", "logprob": 0.0 } ] }, { "token": "unknown", "logprob": -0.00182, "top_logprobs": [ { "token": "unknown", "logprob": -0.0018 } ] }, { "token": "\"}", "logprob": 0.0, "top_logprobs": [ { "token": "\"}", "logprob": 0.0 } ] } ]} A log-probability of −0.00182 is a probability of 99.82%. We asked for five alternatives and got none: Luna put essentially nothing on "true" or "false". And it's wrong. Alan is kind with rough skin, so he's red; red means green; young, round and green means blue. "Alan is not blue" is false, three steps in.
When Luna says 99% One example proves nothing, so we asked for log-probabilities on all 3,600 problems with exactly the benchmark's request. It cost about 14 cents, and every request and response is kept verbatim. We preregistered four predictions first. All four held. Each point is a band of stated probability; the dashed diagonal is where a perfectly calibrated model would sit. Both tasks pooled. Luna stated 99% or more on 2,672 of the 3,600 answers, and 68% of those were right. Below that it's flat: whether Luna says 70%, 90% or 97%, it's right about half the time. Jev stated 99% or more on 1,667 answers and got 98.9% of them right. Its line runs close to the diagonal the whole way. The other numbers say the same thing: Wrong answers stated at 95% or more: 78% of Luna's, against 7% and 17% of Jev's on the two tasks. Expected calibration error, the average gap between stated and actual: 0.32 for Luna on both tasks; 0.04 and 0.03 for Jev.
Separating right from wrong (AUROC, where 0.5 means the probability tells you nothing): 0.66 and 0.68 for Luna; 0.85 and 0.86 for Jev.