Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,686 words · 1 segments analyzed
When I trained the previous ai comment classifier, I used partially personal private data to do it, and built it on a somewhat shaky foundation, so I couldn’t share the code or data. I rebuilt it on public data and a better foundation! First off, you might want to try it out. Nothing you paste into that web page leaves your browser, so you can safely try it with whatever you like. I have invited some testers to try out an earlier version of it, and they had mainly positive feedback to give. We won’t break down robot-isms the way we broke down Claude-isms in the previous article, because in the ui of the new classifier you can just click any part of the text being classified to see which features activate on that portion of the text, and how they contribute to the overall judgment. Here’s an example of the expanded feature activation view. In terms of performance, the headline number is the balanced accuracy of 77 %. This is how often the classifier gets the human vs. robot verdict right, assuming human-written and robot-generated comments are equally likely. The classifier also prints a predicted percentage which is calibrated, meaning it can be read as the probability that any specific verdict is correct. We test this through the calibration curve, which shows what probability the classifier assigns to an event with a known probability. Since all dots lie very close to the reference diagonal, we know they are approximately correct. This holds true across comments of multiple lengths, where a fitted temperature parameter adjusts for increased confidence as the amount of data increases. We can get more details about the classifier’s failure modes by looking at its confusion matrix. In this table, “robot” is considered the positive class, i.e. the thing we want to detect. The abbreviations stand for true/false positive/negative rate. verdict: human verdict: robot input: human tnr = 0.73 fpr = 0.27 input: robot fnr = 0.20 tpr = 0.80 When presented with a known-human input, the classifier correctly judges it as human 73 % of the time. With a known-robot input, it is correctly judged 80 % of the time. This means in both cases (known-human and known-robot) the mistake rate is around 25 %. That might sound high! But remember that this mistake rate is the aggregate over all possible inputs. We don’t need to pay too much attention to it, because the classifier outputs a calibrated predictive percentage every time it classifies something. Thus, for individual judgments, we know when the risk of false positives is lower or higher. When the classifier is very confident – e.g. when the confidence is 80 % or more – the risk of a false positive drops to 5 %. When the classifier is uncertain – when confidence is around 50 % – then by calibration it will issue the wrong verdict around half the time. I mention the numbers in this confusion matrix only because they are so often used when discussing classifiers, so more academically inclined readers may expect to see it. Here are some other requested numbers: Accuracy 77 % Precision 75 % Recall 80 % Sensitivity 80 % Specificity 73 % F1 score 77 % The accuracy, precision, and F1 score depend on the base rate, but here they are computed from an ignorance assumption, i.e. an equal mix of human-written and robot-generated comments. All of these numbers come from cross-validation. I have also manually tested a smaller non-synthetic set of real-world comments from humans and robots to see how well the classifier generalises slightly out of sample. verdict: human verdict: robot input: human tnr = 0.89 fpr = 0.11 input: robot fnr = 0.14 tpr = 0.86 This translates to the following performance numbers: Accuracy 88 % Precision 89 % Recall 86 % Sensitivity 86 % Specificity 89 % F1 score 87 % This is very good! It looks like non-synthetic, more real-worldy cases are easier for the classifier to discriminate between than the training data. Of course, all of this is tested with code comments only. The classifier is not built to detect robot-generated texts of other kinds. It can do it, but I make no promises of its accuracy. With that out of the way, let’s talk about how it’s made. Data collection The first step, as before, is to build a good data set. Ideally, we’d plan this meticulously and do it right the first time. If we do that, it should cost us about $30 to get the dataset that powers this classifier. It contains enough data to reach diminishing returns in discriminating between the more similar models.11 It is possible to extract a more powerful classifier with more data, but it would start to be very expensive since classifier power appears to scale with the log of money spent. That is, if you plan it out and do it right the first time. I didn’t do that. I discovered much later, when evaluating features, that the data I had was junk and I had to collect it all over22 💸. Then after a while I discovered again that the data was still junk and had to be recollected again33 💸💸💸. The general idea was to find a set of permissively licenced or copy-left repositories, check out their latest commit from the year 2021, and then take a few random files from that commit. These contain human comments. Then we strip out all comments from those files, and have llms generate new comments for the same files. That provides us with robot comments. As long as we try to keep the number tokens for each file balanced between all classes (humans and llm models), we can avoid subject matter leakage, where the classifier learns to distinguish files or repositories rather than the style of the text itself. The general idea is simple! But the devil’s where the devil usually is. Here are some mistakes I made, in no particular order: Accidentally picking different source files for each llm to generate comments for. This causes subject matter leakage. Generating llm comments for files with very few human comments. This also causes subject matter leakage over the human–robot barrier. Failing to strip out docstrings when blinding llms to human comments in source files. This causes llms to generate comments more similar to humans because they try to match the existing repository style. Though it should be said this had a smaller effect than I thought it would. Related to the above, some languages support many different syntaxes for comments, and some are used more often than others. Failure to detect existing comments in all syntaxes leaves comments behind to contaminate llm generation, and also makes it hard to get all the data that has been produced. Not filtering out human comments that are very short. Most human comments only say things like “main task structure” or “chIcon” or “Alias” and including those teaches the classifier that humans write like shit. Since llms were instructed to write more detailed comments, it seems reasonable to compare those to more detailed human comments too. Using a fixed prompt for generating llm comments. This results in a dataset with narrower variation than desirable for learning all the quirks needed to separate models and humans. Not all of these problems required regenerating data from scratch. Some could be worked around by filtering and preprocessing the data that already existed. Either way, this was the least fun part of the project, and it cost significantly more than the theoretical $30. Feature evaluation After collecting data, we need to design a classifier that works on that data. This means evaluating candidate features. Doing so isn’t expensive in money, but in cpu time. Evaluating features, in the most powerful sense, means training the classifier on all subsets of candidate features and seeing which performs best. That’s unreasonable, as even with only 15 candidate features, it requires training over 30,000 different classifiers, which need to be trained five ways each for cross-validation to boot. What I ended up doing was guiding the feature selection by the accuracy of classifiers as trained on individual features, for different discrimination tasks. In other words, I had a script that checked “does character frequencies discriminate better between robots and humans than word lengths?” and then repeated that for comparisons between different features, and different classes.44 Different classes means the question is asked not just for robots-vs.-humans but also Claude-vs.-Grok, and GPT-vs.-Gemini, etc. Each of the class pair comparisons produced a list of feature rankings. These lists mostly agreed on the order of features, but there were some disagreements. The ranking of features by power, and the strength of disagreement around relative rankings, is rendered in the graph below. I think the graph reads quite intuitively, but just to be sure: A black arrow means all comparisons agreed on the relative strength of the two features connected with the arrow.55 I suppose technically it means that if a comparison didn’t agree, at least it didn’t disagree. In other words, if four comparisons indicate that feature A and B have roughly the same power, but a fifth comparison indicates feature A is better than feature B, then the graph will show a black arrow from B to A despite the lukewarm response from four out of five comparisons. A blue arrow means at least two comparisons agreed on the relative strength of the two features connected with the arrow, and only one comparison disagreed. A red arrow means more disagreement (and the exact numbers are shown in the arrow label), but the arrow still points in the direction of the dominant opinion. Each feature box also has an information quantity expressed in bits. That shows how much that feature helps, on average, in distinguishing between two classes. Not only is this graph extremely fun to look at – it is also very informative! The feature names may be nonsensical, so we’ll have a brief description of each. As the running example, I will use the following excerpt from a Donald Trump speech: