Pangram verdict · v3.3
We believe that this entire text is AI.
AI likelihood · overall
AIArticle text · 1,379 words · 1 segments analyzed
From pairwise reward modeling to calibrated, multiway decisions Jev looks mysterious when viewed as an alternative to a language model. It becomes much simpler when viewed as the next step in reward modeling. The core idea is: \[ \text{RLCD} = \text{multiway preference modeling} + \text{probability calibration} \] More specifically, RLCD is a schema-conditioned Plackett–Luce objective. Jev turns that objective into a product by adding typed outputs and parallel inference. That is the secret: the reward model is no longer hidden behind a generator. The reward model becomes the model. Figure 1. The learned object changes at each step: a scalar reward becomes a preference, the preference becomes a multiway distribution, and calibration turns that distribution into a decision interface. Reward Modeling Started with a Scalar A conventional reward model receives a context \(x\) and a candidate answer \(a\), then produces a scalar: \[ r_\theta(x,a)\in\mathbb{R} \] Outcome reward models score the final answer. Process reward models score individual reasoning steps. In both cases, the learned object is an absolute-looking number. The problem is that this number is not actually absolute. A reward of \(0.8\) does not have a stable meaning across problems, candidate pools, checkpoints, or model families. It is mainly useful for comparing candidates generated under similar conditions: \[ r_\theta(x,a_1) > r_\theta(x,a_2) \] The operational signal was always relative preference. The scalar merely hid it. PPRM Made the Preference Explicit LLaMA-Berry’s Pairwise Preference Reward Model, or PPRM, exposes the comparison directly. Given a problem \(x\) and two solutions \(a_1\) and \(a_2\), PPRM answers: Is the first answer better than the second answer? Its probability has the form: \[ P(a_1 \succ a_2\mid x) = \frac{\exp u_\theta(x,a_1)} {\exp u_\theta(x,a_1)+\exp u_\theta(x,a_2)} \] Equivalently: \[ P(a_1 \succ a_2\mid x) = \sigma\left( u_\theta(x,a_1)-u_\theta(x,a_2) \right) \] This is the Bradley–Terry model. LLaMA-Berry implements the comparison as a constrained language-model decision over Yes and No tokens. It trains the evaluator on almost 7.8 million mathematical-solution pairs and uses DPO to improve the pairwise prediction task. The essential change is conceptual: reward modeling becomes preference-probability modeling. See the LLaMA-Berry paper. PPRM still contains a latent scalar utility \(u_\theta(x,a)\), but that utility is no longer presented as an absolute reward. It becomes meaningful through a normalized comparison. LLaMA-Berry subsequently uses Enhanced Borda Count to aggregate pairwise comparisons inside MCTS. That is downstream search machinery. EBC neither defines PPRM’s preference loss nor provides the bridge from PPRM to RLCD. The relevant lineage is simply: \[ \text{scalar reward} \rightarrow \text{pairwise preference} \rightarrow \text{multiway preference} \rightarrow \text{calibrated decision} \] Plackett–Luce Is the Multiway PPRM PPRM compares two candidates. A real decision interface usually receives more than two. Let the candidate set be: \[ A=\{a_1,a_2,\dots,a_K\} \] Assign each candidate a context-dependent utility: \[ u_i=u_\theta(x,a_i) \] Then normalize all candidates together: \[ P(a_i\mid x,A) = \frac{\exp u_i} {\sum_{j=1}^{K}\exp u_j} \] This is the Luce choice model, also known as multinomial logit. It is the top-one form of the Plackett–Luce family. When \(K=2\), it reduces exactly to Bradley–Terry: \[ P(a_1\mid x,\{a_1,a_2\}) = \frac{\exp u_1}{\exp u_1+\exp u_2} \] PPRM is therefore the binary case of the same choice geometry. If the supervision contains a complete ranking \[ a_{\pi_1}\succ a_{\pi_2}\succ\dots\succ a_{\pi_K}, \] the full Plackett–Luce likelihood repeatedly selects the next-best remaining candidate: \[ P(\pi\mid x) = \prod_{t=1}^{K} \frac{\exp u_{\pi_t}} {\sum_{j=t}^{K}\exp u_{\pi_j}} \] The corresponding loss is: \[ \mathcal{L}_{\mathrm{PL}} = -\sum_{t=1}^{K} \log \frac{\exp u_{\pi_t}} {\sum_{j=t}^{K}\exp u_{\pi_j}} \] When the label specifies only one correct choice \(y\), the loss becomes: \[ \mathcal{L}_{\mathrm{choice}} = -\log \frac{\exp u_y} {\sum_j\exp u_j} \] That is the first stage of the Plackett–Luce likelihood: a multiway extension of PPRM. This is the mathematical center of RLCD. Figure 2. Bradley–Terry and PPRM are the two-candidate case of the same Luce choice geometry. Plackett–Luce extends that normalization from one choice to a complete or partial ranking. RLCD Adds Calibration Plackett–Luce gives us a probability distribution, but normalization is not calibration. A softmax vector always sums to one. That does not mean a prediction reported as \(0.8\) is correct 80% of the time. Calibration adds that empirical meaning: \[ P(Y=\hat{Y}\mid \hat{P}=p)\approx p \] Across predictions assigned probability \(0.8\), approximately 80% should be correct. This is also the contract TypeSafe gives for RLCD: Jev returns decisions and probabilities, and higher reported probabilities should correspond to higher observed accuracy. See TypeSafe’s RLCD primer. A minimal implementation uses a proper scoring rule such as log loss: \[ \mathcal{L}_{\mathrm{NLL}}=-\log p_y \] Brier calibration: confidence gets a price The Brier score makes the calibration objective concrete. For a binary Noul decision, let \(p=P(Y=1\mid x)\) and \(y\in\{0,1\}\). The score is: \[ \operatorname{BS}(p,y)=(p-y)^2 \] If the model reports \(p=0.8\), it receives a score of \(0.04\) when the event occurs and \(0.64\) when it does not. The confidently wrong forecast costs sixteen times as much as the confidently correct one. This is why the Brier score fits a decision model. It is a strictly proper scoring rule: in expectation, the model minimizes the score by reporting the true conditional probability instead of gaming the threshold. The score was introduced for probabilistic forecasts by Glenn Brier; its role as a proper scoring rule is developed by Gneiting and Raftery. For a multiway Choice, the score extends to the full probability vector. Using the normalization that makes the two-class case match the binary formula: \[ \operatorname{BS}(\mathbf{p},y) = \frac{1}{2} \sum_{i=1}^{K} \left(p_i-\mathbb{1}[i=y]\right)^2 \] This matters because top-1 accuracy discards probability quality. Two models can choose the same action while reporting \(0.55\) and \(0.99\). Once outcomes arrive, Brier score tells us whether that extra confidence was earned. For binary outcomes, the Murphy decomposition separates the mean score into three terms: \[ \operatorname{BS} = \operatorname{REL} - \operatorname{RES} + \operatorname{UNC} \] Reliability \(\operatorname{REL}\) measures the gap between reported probabilities and observed frequencies. Lower is better. Resolution \(\operatorname{RES}\) measures whether the model separates cases with different outcome rates. Higher is better. Uncertainty \(\operatorname{UNC}\) is the base-rate difficulty of the evaluation set. It is fixed when models are compared on the same data. A lower Brier score can therefore come from better calibration, better separation of easy and hard cases, or both. A constant base-rate predictor can be calibrated while having zero resolution; Brier exposes that weakness. Figure 3. Brier score prices confidence, decomposes forecast quality, and closes the loop from observed outcomes to an operational decision policy. An RLCD implementation can apply Brier score to the decision probabilities during training and use it again as a held-out objective for post-hoc calibration. With temperature scaling, the calibration parameter can be selected directly on validation outcomes: \[ T^* = \arg\min_{T>0} \sum_{n=1}^{N} \operatorname{BS}\!\left(\mathbf{p}^{(T)}(x_n),y_n\right) \] Temperature scaling then adjusts the sharpness of the distribution: \[ p_i = \frac{\exp(u_i/T)} {\sum_j\exp(u_j/T)} \] Here \(T\) controls how concentrated the probabilities are without changing their ordering. Brier is the objective; temperature scaling is the calibrator. One measures probability quality, while the other changes the distribution. This separates two objectives that ordinary reward modeling often conflates: Ranking asks whether the best candidate appears first. Calibration asks whether the model knows how often that decision is right. Automation needs both. Ranking selects an action; calibration determines whether software should execute it, defer it, or escalate it. The useful abstraction is: \[ \text{RLCD} = \text{Plackett–Luce preference loss} + \text{calibration constraint} \] Figure 4. Calibration attaches empirical meaning to confidence, allowing application-specific policies to decide when to gather context, escalate, or execute. The reliability curve is conceptual, not a Jev benchmark. Jev Turns the Reward Model into the Product In the conventional RLHF stack, the reward model is an internal component: \[ \text{prompt} \rightarrow \text{generator} \rightarrow \text{candidate response} \rightarrow \text{reward model} \] Users interact with the generator. The reward model only trains or evaluates it. Jev reverses that architecture: \[ \text{state} + \text{candidate schema} \rightarrow \text{calibrated decision distribution} \] There is no need to generate an explanation and parse it back into an action. The evaluator itself becomes the runtime interface. Jev exposes three primitives: Jev primitive Preference-model interpretation Noul Binary Bradley–Terry decision between true and false Choice Luce distribution over \(K\) unordered alternatives Score Distribution over an ordered set of levels A Choice returns the selected option, the complete probability distribution, and a confidence value. A Score returns a position along user-defined levels together with the distribution across those levels. A Noul returns the probability that a proposition is true. See Jev’s primitive documentation.