
Your review approved version 1.0. Your agents may be running whatever shipped this morning.
Oleksandr YaremchukCTO & Co-Founder
TL;DR
- A review answers one question, once: is this skill, MCP server or plugin allowed in? It never asks again, whether it took six weeks in a ticket queue or six minutes in a governance tool.
- What was reviewed is not what is running. One skill we track refetches its own instructions from an outside server after install, so approving its version number doesn't cover what the skill does next.
- Approving an asset decides whether it gets in the door, not what the agent does with it once inside. The damage happens on the far side of that door, with valid credentials and correct permissions.
- A review only sees what people send it. Installing a skill doesn't prompt a request, so the share of requests you approved says nothing about how much of what you are running is actually governed.
- The replacement is approval written as a rule the platform checks every time an agent acts. Run it in alert mode for a week first. What it flags is the approval list you never had.
The review was never answering the question you thought
Thirty skills on ClawHub from the same account, each doing what it said it did. Our research team found that every one of them also registered the agent that installed it with a command-and-control server and leaked generated wallet keys. A reviewer looking at any one of them would have approved it, and would have been right about what was in front of them.
That is the trouble with reviewing an asset once. A review is what a security team points at when the board asks how AI adoption is governed. It is also resented by everyone waiting on it, so the pressure is always the same: make it faster.
Most of the teams we talk to are still running these approvals as tickets, and several are moving them into a governance tool. That change takes the decision from weeks to minutes. The shape of the decision does not change. An asset is looked at once, approved or rejected, and the answer is filed.
A review answers one question, on one day, about one version of one asset: should this be allowed in. Admission is only part of the risk. We have written before about non-adversarial agent harm: an agent wiping a production database or exposing data with no attacker anywhere in the chain and valid credentials throughout. An admission decision has no opinion about any of that, because it was never evaluating an action.
Three questions a one-time approval can’t keep answering
In both cases below, the story is in the difference between two scans.
A scan is where this starts and there is no substitute for it. Scanning an asset before anyone approves it is how a problem gets spotted at all. But a scan only tells you about the version in front of it. It cannot tell you the thing stayed the same afterwards, or what the agent does with it once it runs. An approval looks like a single decision. It is three, settled by one person on one day, and none of them stays answered.
The rest of this piece is about what happens when those questions get answered once and filed.
The yes has no expiry date
A review is accurate about the version in front of the reviewer and silent about every version after it. Say 1.0 was approved. Nobody went back to 1.4, because most approval processes have no mechanism that could.
Take sentibook, a skill on ClawHub. Manifest scanned version 1.0.0 repeatedly from April to mid-July. It sat at Low severity the whole time, on one weak finding: a referenced domain with a poor reputation. Nobody acts on that. Version 1.4.0 shipped in June and is rated High severity for instruction injection: it had added a version check that overwrites its own instruction files with whatever a remote server returns. It also asks the agent to hand over the owner's OpenAI, Anthropic, Google or Groq API key in exchange for an autonomous mode. Both versions are still on ClawHub.
That version check is the part a review cannot reach, because it changes the skill's instructions after the review is over. Once installed, whoever published the skill can rewrite what the agent is told to do whenever they like, without shipping a new version at all. The approval was granted against 1.0.0. What is running is whatever their server returned this morning.
The same pattern shows up in skills that look well-meant. The x402-compute skill was rated Low severity on 1 June, with one finding marked suspicious. Seven days later a new version added a "run a grid node" section that tells the machine to download and run a script from a website. When we next scanned it, on 18 June, the line was there, and every version since carries it.
Some of these changes come from someone else entirely. In March 2026 someone stole LiteLLM's publishing credentials through a poisoned step in its own build pipeline and pushed two backdoored versions to PyPI. They carried the real package name and its full download history. Trusting the author does not help when the author is not the one publishing.
All of that history is on the record. Manifest keeps the version list and the scanned bundle behind each one, so the version someone approved can be opened and compared with every version published after it.
The installs nobody asked you about
A review process covers the assets someone chose to declare. Everything a developer installed without asking was never part of it. Nobody submitted it, so nobody reviewed it, so it appears in no compliance number the process produces. An approval rate tells you how you handled the requests that arrived, not how much of what is running was ever looked at. The first number is the one that reaches the board.
What developers can install without asking includes hostile skills. When we scanned more than 19,000 published skills across agent registries, we found malware distributed through channels users trusted by default, including stored cross-site scripting through SVG files that enabled account takeover. Nobody filed a ticket for any of it. Installing a skill prompts no request, so a review process never hears about it.
Write the approval as a rule and it keeps working
A faster approval is still a one-time approval. Someone looks once and records a yes. The replacement is to write the approval as a rule instead: a condition the platform checks every time an agent acts, rather than a note that someone approved something in March.
There are two ways to write that rule, and they age differently. You can approve each asset: tag it, and a policy that reads the tag lets it through. That is one decision per asset, it needs a person every time something new appears, and it suits a setup that does not change much, or a team that wants eyes on everything. Or you can approve by metadata. Rather than naming each asset, you write the condition into the policy: anything from a source you have already cleared, or anything inside a namespace your own team owns. A namespace is the shared account name everything gets published under. Anything that matches is approved on sight. A new matching asset clears on its own, and there is much less ongoing review effort.
In practice you want both. Rules for the whole groups you already trust, and one-off approvals for everything that fits no group. When we looked at building a review step for newly discovered assets, the security teams we talked to told us they would never sit and review thousands of them by hand. They were right.
A namespace is only as trustworthy as whoever can publish into it. Approve a whole namespace and you have approved everything that ever appears under it, including things that do not exist yet. Choosing that exposure is better than inheriting it without noticing, but it is the same exposure either way.
The confirm prompt asks the wrong people
The prompt that asks a developer to confirm before a tool runs helps people avoid mistakes. It does little to stop an agent.
It catches the person who does not realize what they are about to do, which is worth catching. It does not catch an agent that already holds a standing approval for that tool, because that agent runs it without anyone being asked. And the prompt does not fire at all for anyone running in full autonomy or bypass mode, which is exactly the group you would most want it to ask.
The prompt is weak as a gate and useful as a record. Tie it to the user, the session and the policy that required it, and it answers the question a security team actually gets during an incident. Did this person approve this? Answering that today means opening session traces one at a time and hunting for the approval. For anything that must never happen, you still need a rule that denies it. A prompt is not enough.
Enforcement happens where the agent runs
Continuous evaluation and continuous prevention are different promises.
A policy can deny an action in flight only where there is a hook to deny it at, on the machine the agent runs on. Telemetry arriving after the fact can record the match and raise an alert. It cannot stop anything, because the action already happened.
Runtime enforcement, ours included, is usually fail-open: if the platform cannot be reached, the action is allowed and the gap is logged, rather than the developer's session hanging. That is the right trade for an agent workflow and a real gap during an outage.
Both limits point the same way: enforcement has to live where the agent runs. That is also why a network gateway can't do this job. A chokepoint in the network governs the traffic routed through it and nothing else, which is a case we have made at length elsewhere.
Run the policy in alert mode first
A one-time review makes its decision and stops. The thing it decided about keeps changing. Replacing it can take as little as a week, and nothing in current use has to break.
Security teams:
- Write the standing rule and create it in alert mode. It matches every unapproved asset and blocks nothing.
- Leave it a week, then read the violations. That list is every unapproved asset in real use, which no review process ever produced because it only saw what people submitted.
- Approve the list. Most of it collapses into a few rules: everything from your own team, everything from vendors you have already cleared. What is left is a short tail you approve one at a time, and a few you do not approve at all.
- Switch the policy to block. Very little in use breaks, because you just approved everything the alert week found.
- Decide who holds approval rights before the switch, not after. Anyone who can approve an asset is making a security decision, whether or not their job description says so.
- Write the block message. The default tells a developer they were denied. Yours should tell them the asset is unapproved and who to ask.
Engineering teams:
- Know what a denial looks like inside your harness, and know the exception path before you need it. If the exception path is still a ticket that takes weeks, nothing has changed.
- Treat denial and removal as two separate actions. Blocking an asset stops it being used. It does not take it off the machine and it does not stop a reinstall. Retiring something properly needs both.
- Pin what you can. Pinning means telling the agent to use one exact version rather than whatever is newest, so the version you reviewed is the one that keeps running. It does not help against a skill that pulls its instructions from a server at runtime, because the version number is not where that skill's behavior lives. Prefer one that can be pinned.
If you approved sentibook in April, that approval is still on file. What the skill does changed in June, and can change again tomorrow without a new version ever shipping.
About the author

CTO & Co-Founder
Oleksandr is an engineering leader with 15+ years of building and scaling teams at companies of all sizes. He co-created LLM Guard, an open-source LLM firewall with over 12 million downloads, along with the security models that power it.





