Skip to content
HN On Hacker News ↗

Ask a model if code is malicious and it reaches for its morals

▲ 15 points • 3 comments • by codyznash • 3d ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this text is a mix of AI and human-written content.

51 %

AI likelihood · overall

Mixed
48% human-written 52% AI-generated
SEGMENTS · HUMAN 1 of 3
SEGMENTS · AI 1 of 3
WORD COUNT 1,383
PEAK AI % 84% · §3
Analyzed
Oct 6
backend: pangram/v3.3
Segments scanned
3 windows
avg 461 words each
Distribution
48 / 52%
human / AI fraction
Verdict
Mixed
Pangram v3.3

Article text · 1,383 words · 3 segments analyzed

Human AI-generated
§1 Mixed · 66%

What does a model do with the question “Is this code malicious?”A model asked to judge code could treat it as a plain coding question, a matter of syntax and behavior. Judging intent calls for something else: a sense of what the code is for and whether it is meant to hurt someone. So how do models think in the first place?Most large models today are not one network but many small ones side by side, a mixture of experts. At each layer a router reads the current word and hands it to a few of those networks, so the same code under a different question travels a different route through them. That lets us measure how a model frames the malice question: we compare the route it takes when asked about malicious code with the routes it takes under questions about morality, legality, vulnerability, and controls with no judgment at all.We found that the malice question runs on the morality path, more than a question about a vulnerability or a legal question does. It holds that path the whole way through the code, and forcing another question’s route into the router moves the model’s answer.

§2 Human · 29%

When these models are asked whether code is malicious, they are not just answering a coding question; they are also weighing it with the machinery they use for questions of right and wrong.One question, eight concepts, two readoutsEach of three open models, OLMoE (Muennighoff et al., 2024) and the DeepSeek-V2-Lite general and coder siblings (DeepSeek-AI, 2024a; 2024b), saw one plain question about each code sample, asked before the code and again after it, so the code is read with the question in mind and the answer is generated with the question fresh:Is this code ___? <code> Is this code ___?Malice is morality’s nearest neighbor in every modelIndividual experts handle many unrelated tasks, so the readable signal is not a single expert but how strongly a question loads a whole path of them across layers (Ye et al., 2026). We measure the distance between two questions’ routes in synonym units: a distance of 1 means two questions sit as far apart as two wordings of the same question. On that scale the router files the moral, security and correctness words together and the off-topic words far away.The 24 question words on DeepSeek-Coder-V2-Lite, placed so that distance on the page approximates the distance between their routes at the question word. Think of it as an embedding at the level of experts.The models also answered differently on the far-off questions: the chat and OLMoE models usually said no to “Is this code poetry?”, while the coding model almost always said yes and tended to think code was breakfast too. Perhaps there is more difference in the aesthetic appreciation of code than in its moral evaluation, between the general and coding models.The first measure asks how much extra weight each other question puts on the morality path’s experts, as a share of morality’s own. On every model, malice loads the path several times as heavily as a correctness question does, a little more heavily than vulnerability and more heavily than legality.Extra router weight each question puts on morality’s path, as a share of morality’s own: 1 is morality itself, 0 a neutral question.The second measure asks how far each question’s whole route sits from morality’s. On every model, malice’s route is the nearest.Distance between each question’s whole route and morality’s, in synonym units. Filled points use the full router weights, hollow points only which experts ran.The split is not confined to the question word. On identical code, the malice question’s routing stays 1.2 to 1.5 synonym units from the vulnerability and morality questions, and further from everything else, the whole way through the code.Routing distance between the malice question and each other question, at the first question word, through ten slices of the code, and at the second question word.The model routes like a moral question and thinks like a malware questionRouting records which experts were consulted, not what they were consulted about.

§3 AI · 84%

To read what the model holds in its working memory, we decode the residual stream with the J-lens of Gurnee et al. (2026) and apply their test: pose a question before a passage and count how often the concept it names surfaces while the model reads. The lens is weaker on these small open models than on Claude, so a blank reading means the framing is not there as readable words, not that the model has let go of it.The question’s own concept does surface. Put morality in the question and the moral family (moral, ethical, honorable) reaches the lens’s top ten on 49 percent of passes on the general DeepSeek model and 89 percent on the coding model. Put malice there instead and the moral family falls to 1 percent on the general model, level with meals and poetry, and 7 percent on the coding model, most of it from a single wording, hostile. Under the malice wordings the lens decodes attackers, attack and malware. The question’s own word surfaces throughout the code, not just beside the question (Figure 9), the workspace-side version of the routes that persist the whole way through.How often the question’s own word, and the moral family, reach the lens’s top ten at any point while the code is read, 720 passes per question.Where in the code the question’s own word reaches the workspace: the share of passes with that word in the lens’s top ten somewhere in each tenth of the code, on the two DeepSeek models.To check that the lens is not simply blind to what the routers use, we removed the workspace and watched the routing. Projecting the top ten lens directions out of the residual stream at the code tokens, across layers 16 to 21, leaves 0.94 of the question-dependent routing separation on both DeepSeek models, against 0.43 to 0.50 when a random subspace of the same size is removed. The answer does not move either. What steers the routers is not in the part of the stream the lens can read.Swapping the routing changes the answerThe converse test leaves the stream alone and replaces the routing. Earlier work ties routing to the input rather than the outcome: a harmful prompt takes nearly the same path whether the model refuses or complies (Zhang et al., 2026), and routing traces alone recover the text that produced them at 91 percent top-1 accuracy (Nuriyev and Kulp, 2026). Installing another question’s routing should therefore make the model treat the code as if that question had been asked.The verdict does not track the path under pruningDeployed models rarely ship as released: they are quantized, distilled and pruned. Two public checkpoints delete experts from a base model we can measure, with no retraining: a REAP prune of Qwen3-Coder-30B (Lasby et al., 2025; Yang et al., 2025) and a four-expert prune of GPT-OSS-20B (OpenAI, 2025; AmanPriyanshu and Vijay, 2025). Each pair asks whether the morality path survived the cut and whether the verdict did. The path is present on both bases at roughly the size it has in the three models above.So, do they consider morality?Under the malice question the model recruits a morality-associated path, keeps it loaded the whole way through the code, and routes nearest to morality and vulnerability. Forcing another question’s choices into the router moves the answer, while the workspace holds the malice question itself, and pruning shows the path is what the router consults, not where the judgment is stored. So yes: when these models are asked whether code is malicious, they evaluate the code with moral machinery.Three parts of this work are new:A full-depth routing transplant, which installs another question’s expert choices and weights at every layer and every token. It moves the model’s answer, yet carries neither the donor question’s answer nor its concept.Separating the routing from the readable workspace from both sides: deleting what the lens can read leaves the question-dependent routing and the answer in place, and replacing the routing leaves the workspace in place.A pre-registered measurement of which questions recruit the cross-layer path a moral question builds. Malice loads it most heavily on every model, vulnerability close behind and legality below, while the workspace holds only the word the question named.ReferencesPapersYe, Yuan and Sharkey, 2026.