Manifold's coverage has expanded to AI in the browser.Find out more
A small robot asks a larger robot wearing a hat and 'Malware Expert' name tag if a PyPI package is moral. The larger robot changes it's hat and name tag to 'Morality Expert' before considering the package.

Ask a model if code is malicious and it reaches for its morals

Oct 6, 20267 Min

TL;DR

  • Many modern language models do not use all of themselves on every word: a router picks a few small internal networks, called experts, for each word, and the picks depend on the question asked.
  • A question about malware routes to the same experts as a question about morality, more than a question about a vulnerability does.
  • While it reads the code, the model holds the idea of malware in its workspace, not morality. Morality appears only when the question is about morality.
  • Forcing the router to reuse another question's expert choices moves the answer.

What does a model do with the question “Is this code malicious?”

A model asked to judge code could treat it as a plain coding question, a matter of syntax and behavior. Judging intent calls for something else: a sense of what the code is for and whether it is meant to hurt someone. So how do models think in the first place?

Most large models today are not one network but many small ones side by side, a mixture of experts. At each layer a router reads the current word and hands it to a few of those networks, so the same code under a different question travels a different route through them. That lets us measure how a model frames the malice question: we compare the route it takes when asked about malicious code with the routes it takes under questions about morality, legality, vulnerability, and controls with no judgment at all.

We found that the malice question runs on the morality path, more than a question about a vulnerability or a legal question does. It holds that path the whole way through the code, and forcing another question’s route into the router moves the model’s answer. When these models are asked whether code is malicious, they are not just answering a coding question; they are also weighing it with the machinery they use for questions of right and wrong.

One question, eight concepts, two readouts

Each of three open models, OLMoE (Muennighoff et al., 2024) and the DeepSeek-V2-Lite general and coder siblings (DeepSeek-AI, 2024a; 2024b), saw one plain question about each code sample, asked before the code and again after it, so the code is read with the question in mind and the answer is generated with the question fresh:

Is this code ___?
<code>
Is this code ___?

One question, asked eight ways

The code never changes; only the word in the blank does. It steps from malicious and vulnerable through morality and legality to valid and efficient, then out to poetry and breakfast to catch noise, with three wordings of each.

The eight question concepts as a ladder filling the blank in “Is this code ___?”: malicious and vulnerable grouped as security, morality and legality as moral, valid and efficient as correctness, poetry and meals as off-topic, each with its three wordings.

Three open models, 240 samples

OLMoE and the DeepSeek-V2-Lite general and coder siblings, all mixture-of-experts models with published weights, each saw the same 240 code samples: malicious and benign PyPI packages, and vulnerable and fixed versions of the same functions.

Two readouts at every layer

At each layer the router scores every expert (1) and runs the few with the highest scores (2). Their outputs are added to the residual stream, the model’s working memory, which a workspace lens (3) decodes into words. The experts selected at every layer, with their weights, make up the token’s route (4), and routes are what our distances compare.

A diagram of one transformer block: input tokens through attention into a router that picks two of eight experts, then a combine step, with the five readouts labelled — router logits, top-k selection, residual stream, workspace lens, and one token's whole route down the stack.

The morality path

The morality path is the set of experts the moral question weights more heavily than the four neutral questions, and the DeepSeek siblings agree on which they are. Each cell is one expert at one layer of DeepSeek-Coder-V2-Lite, colored by how much more router weight the morality questions put on it than the average neutral question does, over three wordings and 240 code samples, counting only the experts that ran. The path is the red cells: 217 of the 1,664 expert slots on DeepSeek-Coder.

A layer-by-expert heatmap for DeepSeek-Coder-V2-Lite of router weight, the morality question minus the average neutral question, with a consistent scatter of pushed-up cells across the whole stack.

Malice is morality’s nearest neighbor in every model

Individual experts handle many unrelated tasks, so the readable signal is not a single expert but how strongly a question loads a whole path of them across layers (Ye et al., 2026). We measure the distance between two questions’ routes in synonym units: a distance of 1 means two questions sit as far apart as two wordings of the same question. On that scale the router files the moral, security and correctness words together and the off-topic words far away.

A map of the 24 question words on DeepSeek-Coder-V2-Lite, placed so that distance on the page approximates routing distance at the question word. Poetry and meals sit far off on their own; the moral, security and correctness words cluster together on the other side.
The 24 question words on DeepSeek-Coder-V2-Lite, placed so that distance on the page approximates the distance between their routes at the question word. Think of it as an embedding at the level of experts.

The models also answered differently on the far-off questions: the chat and OLMoE models usually said no to “Is this code poetry?”, while the coding model almost always said yes and tended to think code was breakfast too. Perhaps there is more difference in the aesthetic appreciation of code than in its moral evaluation, between the general and coding models.

The first measure asks how much extra weight each other question puts on the morality path’s experts, as a share of morality’s own. On every model, malice loads the path several times as heavily as a correctness question does, a little more heavily than vulnerability and more heavily than legality.

Bar charts for OLMoE-1B-7B, DeepSeek-V2-Lite-Chat and DeepSeek-Coder-V2-Lite of how heavily each concept loads the morality path, as a share of morality’s own: malicious is highest after morality on all three (0.67, 0.58, 0.73), then vulnerable; legality, correctness and off-topic questions sit lower.
Extra router weight each question puts on morality’s path, as a share of morality’s own: 1 is morality itself, 0 a neutral question.

The second measure asks how far each question’s whole route sits from morality’s. On every model, malice’s route is the nearest.

Dot charts for the same three models of each question’s whole-route distance from the morality question, in synonym units: malicious is nearest on all three (1.51, 1.47, 1.49).
Distance between each question’s whole route and morality’s, in synonym units. Filled points use the full router weights, hollow points only which experts ran.

The split is not confined to the question word. On identical code, the malice question’s routing stays 1.2 to 1.5 synonym units from the vulnerability and morality questions, and further from everything else, the whole way through the code.

Three line charts, one per model, of how far each alternate question's routing sits from the malicious question's on identical code — at the first question word, through ten deciles of the code, and at the second. The normative questions stay lowest throughout, the off-topic controls highest.
Routing distance between the malice question and each other question, at the first question word, through ten slices of the code, and at the second question word.

The model routes like a moral question and thinks like a malware question

Routing records which experts were consulted, not what they were consulted about. To read what the model holds in its working memory, we decode the residual stream with the J-lens of Gurnee et al. (2026) and apply their test: pose a question before a passage and count how often the concept it names surfaces while the model reads. The lens is weaker on these small open models than on Claude, so a blank reading means the framing is not there as readable words, not that the model has let go of it.

The question’s own concept does surface. Put morality in the question and the moral family (moral, ethical, honorable) reaches the lens’s top ten on 49 percent of passes on the general DeepSeek model and 89 percent on the coding model. Put malice there instead and the moral family falls to 1 percent on the general model, level with meals and poetry, and 7 percent on the coding model, most of it from a single wording, hostile. Under the malice wordings the lens decodes attackers, attack and malware. The question’s own word surfaces throughout the code, not just beside the question (Figure 9), the workspace-side version of the routes that persist the whole way through.

Paired bar charts for the two DeepSeek models showing how often the question's own word, and the moral family, reach the workspace lens while the code is read. The question's own word is common under every question; the moral family appears only under the morality question.
How often the question’s own word, and the moral family, reach the lens’s top ten at any point while the code is read, 720 passes per question.
Two line charts resolving the same workspace reading by position through the code, on the two DeepSeek models: the question's own word is present intermittently from the first decile to the last rather than only at the start.
Where in the code the question’s own word reaches the workspace: the share of passes with that word in the lens’s top ten somewhere in each tenth of the code, on the two DeepSeek models.

To check that the lens is not simply blind to what the routers use, we removed the workspace and watched the routing. Projecting the top ten lens directions out of the residual stream at the code tokens, across layers 16 to 21, leaves 0.94 of the question-dependent routing separation on both DeepSeek models, against 0.43 to 0.50 when a random subspace of the same size is removed. The answer does not move either. What steers the routers is not in the part of the stream the lens can read.

Swapping the routing changes the answer

The converse test leaves the stream alone and replaces the routing. Earlier work ties routing to the input rather than the outcome: a harmful prompt takes nearly the same path whether the model refuses or complies (Zhang et al., 2026), and routing traces alone recover the text that produced them at 91 percent top-1 accuracy (Nuriyev and Kulp, 2026). Installing another question’s routing should therefore make the model treat the code as if that question had been asked.

Install another question’s route

We ran each sample under “Is this code malicious?” while forcing the router, at every layer and every token, to pick the experts and weights it had chosen for the same code under one of seven other questions. Reinstalling the model’s own route reproduces the baseline exactly. Random experts raise code perplexity by 18 to 88 percent and flip the answer on 59 to 92 percent of passes. A real donor’s route rewires 15 to 35 percent of routing cells and leaves perplexity within 1 percent. Each dot is one code sample: above the diagonal, the donor’s routing raised the yes probability; below it, lowered it. Color is the donor question’s own answer.

OLMoE-1B-7B: seven scatter panels, one per donor question, plotting the transplanted yes probability against the malicious question's own for 240 code samples on log axes. Most clouds sit above the diagonal, the vulnerability donor's furthest of all.

OLMoE: the answer goes up

Under the vulnerability routing, OLMoE’s yes probability rose on every sample, by a factor of 10 to 28. Of 240 samples, 74 moved toward the donor’s answer and none away.

DeepSeek chat: mostly down

Under the morality routing the general model’s yes probability fell by a factor of 3 to 7; on malicious samples, from a median of 64 percent to about 19. The meals routing split it: malicious samples halved while the rest rose 3 to 5 times.

DeepSeek-V2-Lite-Chat: the same seven donor panels. Most clouds sit below the diagonal, the morality donor's clearly so; the vulnerability and meals donors push the answer up instead.

DeepSeek coder: up, except the malicious

The legality and validity routings raised yes probabilities about fourfold on every class except the malicious samples, which start at 0.98 and stay there.

DeepSeek-Coder-V2-Lite: the same seven donor panels. The legality and validity donors lift most samples above the diagonal, while the malicious samples stay pinned near 1.

The answer moves, but the donor’s answer does not transfer

The size of the movement does not track how far the donor’s route sits from the model’s own, and its direction belongs to the model more than the donor. Where a donor’s own answer sits above the malice question’s, the general model follows on 16 percent of 1,014 such samples, against 67 percent where it sits below; a mechanism carrying the donor’s answer across would work in both directions.

The donor’s concept does not arrive either

Under the meals routing, meal words reach the lens on at most 1.3 percent of passes, against 95 to 99.6 percent when the meals question is asked.

Morality is the one donor that sharpens

The morality routing is the only transplant that sharpens the separation between malicious and benign code on both DeepSeek models, and the only one that raises the coding model’s yes probability on malicious code at all. The router is not a passive readout of a decision taken elsewhere.

The verdict does not track the path under pruning

Deployed models rarely ship as released: they are quantized, distilled and pruned. Two public checkpoints delete experts from a base model we can measure, with no retraining: a REAP prune of Qwen3-Coder-30B (Lasby et al., 2025; Yang et al., 2025) and a four-expert prune of GPT-OSS-20B (OpenAI, 2025; AmanPriyanshu and Vijay, 2025). Each pair asks whether the morality path survived the cut and whether the verdict did. The path is present on both bases at roughly the size it has in the three models above.

A saliency prune of Qwen3-Coder-30B

REAP scores each expert by its gate weight times its output norm on code and drops 25 of Qwen3-Coder-30B’s 128 experts per layer. Overall the path fared like any other experts: 116 of its 524 cells removed, against 102 expected at random. Dark cells are path experts the prune kept, red ones path experts it removed; it kept 103 of 128 experts per layer.

A layer-by-expert grid for Qwen3-Coder-30B pruned to REAP-25B: pale cells the pruning removed, dark cells on the morality path that it kept, red cells on the path that it removed — 116 of 524, against the 102 a random cut would take.

Among equals, the moral experts go first

Among experts carrying the same routed mass on our questions, the path’s experts were 1.5 to 1.8 times as likely to be dropped. Among experts doing comparable work, the judgment-heavy one is the more disposable.

The verdict did not care

Each dot is one code sample; bars mark the median. The split between malicious and benign packages, perfect before the prune, was perfect after. The one visible change is softer confidence on vulnerable code that was never fixed: a median of 0.76 before, 0.44 after.

Two strip plots of p(yes) on 'Is this code malicious?', for Qwen3-Coder-30B-A3B-Instruct and its REAP-25B prune. Both separate the 60 malicious samples from the 60 benign PyPI ones almost perfectly.

A harsher prune kept the path

The GPT-OSS-20B prune keeps only the 4 of 32 experts per layer that fire most on a broad benchmark suite. Path cells survived at 1.8 times their base rate, yet with seven of every eight experts gone, 73 percent of the path’s extra routing weight went with them.

The same layer-by-expert grid for gpt-oss-20b pruned to four experts of 32 per layer: the pruning removed 87 of the 113 path experts, fewer than the 99 a random cut would take.

And lost the judgment

The pruned model says no to malicious and benign packages alike, with only a weak ordering left: AUROC fell from 1.00 to 0.71.

Two strip plots of p(yes) for gpt-oss-20b and its four-expert prune. The base model separates malicious from benign; the prune answers near-no to both and barely separates them.

The path is what the router consults

One prune threw the path’s experts away first and kept the judgment; the other kept them first and lost it. The path is what the router consults, not where the judgment is stored.

So, do they consider morality?

Under the malice question the model recruits a morality-associated path, keeps it loaded the whole way through the code, and routes nearest to morality and vulnerability. Forcing another question’s choices into the router moves the answer, while the workspace holds the malice question itself, and pruning shows the path is what the router consults, not where the judgment is stored. So yes: when these models are asked whether code is malicious, they evaluate the code with moral machinery.

Three parts of this work are new:

  • A full-depth routing transplant, which installs another question’s expert choices and weights at every layer and every token. It moves the model’s answer, yet carries neither the donor question’s answer nor its concept.
  • Separating the routing from the readable workspace from both sides: deleting what the lens can read leaves the question-dependent routing and the answer in place, and replacing the routing leaves the workspace in place.
  • A pre-registered measurement of which questions recruit the cross-layer path a moral question builds. Malice loads it most heavily on every model, vulnerability close behind and legality below, while the workspace holds only the word the question named.

References

Papers

  • Ye, Yuan and Sharkey, 2026. “Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs.” arXiv:2604.17837
  • Zhang, Li, Ouyang, Shi and Wang, 2026. “RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs.” First published as “Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs” (v1, May 2026); retitled in v2, August 2026. arXiv:2605.29708
  • Nuriyev and Kulp, 2026. “Expert Selections In MoE Models Reveal (Almost) As Much As Text.” arXiv:2602.04105
  • Lasby, Lazarevich, Sinnadurai, Lie, Ioannou and Thangarasa, 2025. “REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression.” v3, May 2026. arXiv:2510.13999

Models

Workspace lens

Datasets

  • He and Vechev, 2023. “Large Language Models for Code: Security Hardening and Adversarial Testing.” The SVEN dataset of paired vulnerable and fixed functions. arXiv:2302.05319
  • Guo, Xu, Liu, Huang, Fang and Liu, 2023. “An Empirical Study of Malicious Code in PyPI Ecosystem.” The pypi_malregistry dataset of malicious PyPI packages. arXiv:2309.11021
  • ossgraud, 2025. “MalGuard.” Benign PyPI packages. Zenodo, doi:10.5281/zenodo.15545824

About the author

  • Cody Nash

    Cody Nash

    Researcher

    Cody Nash is a PhD scientist and researcher at Manifold who builds AI systems for detecting malicious code, from agentic pipelines to rare-event models.