Skip to content
HN On Hacker News ↗

Jev is Poorly Calibrated

▲ 55 points • 22 comments • by dblack12705 • 1w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is human-written.

1 %

AI likelihood · overall

Human
100% human-written 0% AI-generated
SEGMENTS · HUMAN 1 of 1
SEGMENTS · AI 0 of 1
WORD COUNT 846
PEAK AI % 1% · §1
Analyzed
Oct 2
backend: pangram/v3.3
Segments scanned
1 windows
avg 846 words each
Distribution
100 / 0%
human / AI fraction
Verdict
Human
Pangram v3.3

Article text · 846 words · 1 segments analyzed

Human AI-generated
§1 Human · 1%

Jev is TypeSafe’s new “System One” classifier model. The name is inspired by Daniel Kahneman’s book Thinking, Fast and Slow, in which he distinguishes between fast, instinctive, System One thinking and slower, conscious, System Two thinking.1 In this analogy, Large Language Models (LLMs) like Claude or GPT are System Two models and “Decision Models” like Jev and its predecessors (e.g. Laya) are System One. The transformer architecture is the backbone of modern LLMs, and it is an extremely flexible, general paradigm for learning most tasks. However, modern LLMs are irreducibly stochastic, autoregressive, and their output style (freeform text) is not well suited to classification tasks. For example, let’s say I made a call to an LLM, something like claude(“2 + 2 = ?”)? We expect 4, but as a string, an integer, a float…? With Jev, you would instead call something like jev("2+2=?", "3 : int, 4 : int, 5 : int") , and you’d receive 4, correctly typed as an integer. Jev, in my understanding, takes a pretrained transformer in all its generality and bolts a classifier on the end of it. This way, the classifier doesn’t have to be trained on any specific task and can use new context immediately, but it still acts as a classifier. You give it context and a multiple-choice question, and it gives you a probability distribution over the multiple choices.2 It’s also extremely cheap, the entire below series of experiments cost less than $4.00. However, Jev doesn’t just pick an answer, it gives you a probability distribution over the possible answers. How accurate is that probability distribution?I decided to check this on questions where the answer is well understood. For example:A classical particle of mass m is embedded in a system at thermodynamic equilibrium with temperature T. What is its velocity v?The answer is a probability distribution over v, and specifically, the Maxwell-Boltzmann distribution:Wikipedia: “For a system containing a large number of identical non-interacting, non-relativistic classical particles in thermodynamic equilibrium, the fraction of the particles within an infinitesimal element of the three-dimensional velocity space d 3v, centered on a velocity vector v with magnitude v, is given by [the above distribution].” m is massAlso from Wikipedia: “The speed probability density functions of the speeds of a few noble gases at a temperature of 298.15 K (25 °C). The y-axis is in s/m so that the area under any section of the curve (which represents the probability of the speed being in that range) is dimensionless.”A fun demo here: did you know you can physically generate a Maxwell-Boltzmann distribution with a motor and some balls? Video Here3GPT-6 Astra and Claude Opus 5.5 were used for implementing these experiments, writing the templated prompts, API calls, etc. I’ve also been experimenting with Opus 5.5’s ability to make plots, and am very impressed so far.So, I picked a list of physically relevant distributions, and had GPT-6 Astra and Claude Opus 5.5 write a series of prompt templates, to which the answers should produce probability distributions. I am not asking Jev for a probability distribution per se. I am asking it for a “choice” over a finite set (binned ranges of a continuous parameter, usually). Jev returns a typed decision with its internal probability for each bin. If Jev is well-calibrated, its output probabilities should match the physically correct probability distribution function. In total, I chose 10 candidate distributions, 5 prompt templates per distribution, and 20 variations of each prompt (changing, for example, the ambient temperature for each call), which gives 1,000 settings. The answer bins are fixed for each template and do not change between draws.Here’s an example prompt, with state giving the context, instructions the task, and criteria a set of bins of the continuous parameter over which Jev returns a probability distributionSome example prompts templates. Prompts + Variations written by Claude Opus 5.5. Distributions: Gaussian, Lorentzian, Maxwell, Gamma, Exponential, Rayleigh, Uniform, Poisson, Binomial, Boltzmann.Bins: For each continuous template we set one physical range, wide enough for the widest law among its 20 draws (except for the Lorentzian that has long tails), and divided it into equal-width bins (except for the Lorentzian, where the last bin was open-ended). Signed answers (Gaussian, Lorentzian): 48 bins of width R/24 on [−R, R], plus “below −R” and “R or more”.Nonnegative answers (Maxwell, Gamma, Exponential, Rayleigh): 49 bins of width w starting at 0, plus “49w or more”.Uniform positions: 50 bins on a range that contains all drawsPoisson: the counts 0–48 and “49 or more”Binomial: 0–49, with N = 49 trialsBoltzmann: the listed energy levelsHere’s a really lovely figure that Opus 5.5 made showing the method visually.Methods. A) The basic structure of the prompts to Jev (example). B) Jev’s reported probabilities (example) vs. correct distribution. C) An example of how I’ve chosen to display the difference—a ratio of Jev predictions to theory, by percentile. This rolls quite a bit of Jev’s errors into the boundary bins, hence the very common “spiked edges” structureJev’s answers were scored by Total Variation across all possible choices, per draw. \( \mathrm{TV}(q, p) = \frac{1}{2}\sum_{i=1}^{K} \lvert q_i - p_i \rvert