Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
MixedArticle text · 972 words · 2 segments analyzed
JevBench v1.3.0JevBench v1.3.0 · our own benchmarkJev-class modelsJevBench is Benchmark Heaven's own benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out.Version 1.3.0 measures 52 systems on the unchanged 534 decisions, including 220 hard ones, and ranks them by the JevBench Score. Built and run by us, not collected from someone else's leaderboard; the results describe the tested configurations, not every application.Scored 21 Sept 2026 · protocol jevbench::v1.2 · 72 easy + 96 standard + 146 judge + 220 hard decisions · one request at a time from a server in Germany · harness, public tasks & scoring rules (MIT) · results JSON sha256 20fce8e6e4f0… · v1.0 resultsJevBench v1.3.0 · 534 decisions per systemJevBench Score (Intelligence, Calibration, Speed, Cost — 25 % each)OfficialIntelligence above chance, Calibration, Speed, Cost — 25 % each, geometric mean; below 50 Intelligence receives a growing near-chance penalty.
Change the weighting ↓1Jev 1.13.074.4I 86 · C 83 · S 83 · K 52 · $0.0402SemIf (Qwen3.5-4B)73.1I 79 · C 73 · S 84 · K 59 · ~$0.022 est.3djev (Maisa, diffusion-gemma)†73.0I 83 · C 65 · S 91 · K 58 · $0.026 ann.4Winnow-12B Q8†71.2I 82 · C 72 · S 82 · K 53 · ~$0.037 est.5reflex 4B†70.3I 80 · C 75 · S 68 · K 60 · ~$0.022 est.6jqv†68.6I 79 · C 79 · S 75 · K 47 · ~$0.056 est.7decision-machine-1†68.3I 62 · C 70 · S 93 · K 54 · $0.0358decider-35b-a3b†67.6I 80 · C 72 · S 81 · K 45 · ~$0.067 est.9open-alternative-jev (Qwen3.5-4B)†67.0I 64 · C 63 · S 83 · K 60 · ~$0.022 est.10system-one-open66.6I 70 · C 57 · S 77 · K 65 · ~$0.015 est.11OpenJev (razorback16)66.4I 79 · C 65 · S 83 · K 45 · ~$0.066 est.12SimpleJev Qwen3.8-27B†66.3I 85 · C 81 · S 71 · K 39 · ~$0.104 est.13ZeroEntropy zerank-2†66.0I 63 · C 76 · S 79 · K 50 · $0.04714GPT-5.6 Luna (low)65.9I 95 · C 90 · S 78 · K 28 · $0.24215openjev-sglang65.3I 83 · C 77 · S 77 · K 36 · ~$0.131 est.16Qwen3-Reranker-4B†63.8I 64 · C 67 · S 79 · K 49 · $0.05017reflex-27b†63.3I 86 · C 86 · S 67 · K 32 · ~$0.181 est.18LitJev†62.7I 82 · C 84 · S 67 · K 34 · ~$0.163 est.19kev 0.6B†62.5I 52 · C 51 · S 76 · K 76 · ~$0.0063 est.20SimpleJev Qwen3.6-35B-A3B†62.5I 80 · C 67 · S 75 · K 38 · ~$0.116 est.21djev†62.4I 81 · C 93 · S 75 · K 27 · ~$0.274 est.22jev-local†61.8I 71 · C 69 · S 69 · K 43 · ~$0.077 est.23decider-2b†61.7I 61 · C 47 · S 83 · K 61 · ~$0.020 est.24Bespoke Nimble 9B†60.5I 78 · C 65 · S 79 · K 33 · ~$0.166 est.25Gemini 3.1 Flash-Lite60.1I 86 · C 68 · S 82 · K 27 · $0.26426OpenJev†60.0I 88 · C 70 · S 76 · K 28 · ~$0.255 est.27kev 4B†59.7I 65 · C 42 · S 76 · K 62 · ~$0.019 est.28DeepSeek V4.1 Flash57.5I 94 · C 97 · S 72 · K 17 · $0.59429kev 8B†56.4I 69 · C 44 · S 75 · K 44 · ~$0.073 est.30Open-Jev 9B†55.0I 71 · C 63 · S 72 · K 28 · ~$0.249 est.31system-one54.8I 70 · C 37 · S 84 · K 41 · ~$0.089 est.32jeff†54.4I 47 · C 65 · S 63 · K 77 · ~$0.0060 est.33Laya†54.4I 46 · C 62 · S 71 · K 86 · ~$0.0029 est.34Open-Jev 2B†51.3I 61 · C 55 · S 73 · K 28 · ~$0.249 est.35OpenDecision†40.6I 41 · C 56 · S 80 · K 75 · ~$0.0066 est.36openJev Verdict 1.4†38.9I 39 · C 74 · S 78 · K 82 · ~$0.0039 est.37openJev Verdict†38.1I 40 · C 51 · S 77 · K 83 · ~$0.0037 est.38kev 0.5B†33.2I 38 · C 47 · S 77 · K 76 · ~$0.0063 est.39GLiNER2 large†29.6I 40 · C 24 · S 62 · K 73 · ~$0.0077 est.40smalljev semantic-v9†27.4I 35 · C 59 · S 80 · K 58 · ~$0.025 est.41GLiNER2†24.0I 36 · C 24 · S 72 · K 83 · ~$0.0037 est.42open-jev-deberta-v3-large23.1I 32 · C 66 · S 66 · K 74 · ~$0.0073 est.43GLiNER2.5 multi†16.6I 28 · C 56 · S 68 · K 82 · ~$0.0039 est.44GLiNER2.5 small†13.8I 26 · C 47 · S 78 · K 82 · ~$0.0039 est.45Mixedbread mxbai-rerank-base-v2†0.8I 7 · C 83 · S 88 · K 68 · $0.01246BAAI bge-reranker-v2-m3†0.7I 6 · C 84 · S 90 · K 73 · $0.007747Alibaba GTE Reranker ModernBERT-base†0.3I 5 · C 77 · S 91 · K 70 · $0.01048Certo v1†0.0I 0 · C 82 · S 94 · K 100 · ~$0.0010 est.classifier.dev (fast tier)† (honorable mention)83.6I 85 · C 78 · S 88 · K 84 · ~$0.0033 est.Qwen3.8 27B (partial run)24.8I 67 · C 92 · S 61 · K 0 · ~$2.669 est.Needle 3, options as tools (partial run)1.1I 14 · C – · S 53 · K 65 · ~$0.014 est.Needle 3 (partial run)0.1I 5 · C – · S 60 · K 59 · ~$0.024 est.Score = Intelligence0.25 × Calibration0.25 × Speed0.25 × Cost0.25 (each 0–100; geometric mean; below 50 Intelligence, × (I / 50)²)Jev (TypeSafe, closed)Jev rebuild (open, or open source planned)Instruction model, JSON schemaSmall tool-calling modelService built on JevZero-shot classifier (not a Jev rebuild)Closed decision model (API only, not Jev)Shown, not ranked — honorable mention (runs another entrant's model) · partial run⏱ Latency of self-hosted and demo endpoints is adjusted ×2 (+0.15 s on our own servers) to approximate production load — an assumption, not a measurement; raw measurements are in the table and the repo.I, C, S, K = Intelligence, Calibration, Speed, Cost; est./ann.