Manifold's coverage has expanded to AI in the browser.Find out more
Illustration of robots at a fashion show. In the front row, a large robot gestures smugly toward the runway while a smaller robot beside him sits with folded arms, unimpressed. Robot models walk the runway, and the caption reads "JEV, SO HOT RIGHT NOW."

Jev, system one models, and how to use them in security

Oct 1, 202613 Min

TL;DR

  • What it is: Jev isn't a smarter LLM. It answers narrow, predefined questions in milliseconds, for almost nothing, and says how sure it is. It has launched a new wave of innovation in the fast, low-cost model space.
  • Where it fits in security: use it as a filter. It clears the cases it’s confident about; everything else goes to rules, a person, or a stronger model.
  • Before you build on it: attackers can steer it through its input, and it only runs as a hosted service. Test it on your own data first, or try an open alternative, which now scores within a few points of Jev.

In the beginning

Jev was the first of a new wave of “system one” models: they think fast and cost little. They are an alternative to the large language models that run their entire process on every token. LLMs are the heavy hammer that has been applied to every problem. But for applications with a high volume of tasks, including checking for prompt injection or monitoring logs of agent behavior, LLMs can be too expensive. Expert systems and classical machine learning covered the easy share of these tasks. This new wave of system one models leverage the generalizability of the transformer architecture to bring more intelligence to the rest.

Security runs on small decisions made at high volume

Is this email phishing? Is this prompt an injection attempt? Should this agent be allowed to make this tool call? Does this alert deserve an analyst? None of these are hard, but all of them happen millions of times.

Until now, those decisions lived in rules and workflows (i.e. Security Orchestration, Automation, and Response (SOAR) playbooks), which break on anything unanticipated; in trained classifiers, which need labeled data for every question; or in LLMs, which are too slow and expensive to run on every event. Jev claims a fourth option: a model's judgment at a classifier's speed and cost. It handles the messy, ambiguous text that breaks a rule, like a phishing email with no known indicators or a prompt phrased to slip past a regex. The workflow still decides what to ask and what to do with the answer. Jev takes over the brittle decision logic inside it; the playbook stays.

Where security decisions live

One event, a phishing email with no known indicators, sent to each of the four places a decision can live.

Phishing email no known indicators

sent to each of

Rules and playbooks
No rule matches the email passes
Trained classifier
No model for this question label data and train first
LLM
A good verdict seconds and cents, every event
Jev
Verdict and confidence at classifier speed and cost
Figure 1: Various approaches to decision making. A phishing email with no known indicators slips past rules, has no trained classifier waiting for it, and costs an LLM seconds and cents to judge. Jev claims to return a verdict with a confidence score at a classifier's speed and cost.

Behind the curtain (what Jev is)

TypeSafe calls Jev its first System One model: it's built for fast, intuitive decisions inside software, not for slow, deliberate reasoning. It doesn't write text. You give it a block of input (the state) and a predefined set of questions, and it answers them all in one parallel call instead of one token at a time. Each answer is a Choice among up to 255 options, a Score on a scale you define, or a Noul, a yes-or-no question returned as a probability (more details here: thoughts.jock.pl, firecrawl).

Three questions, two answer shapes

Top: an LLM returns one string per call. Bottom: Jev returns a probability per declared answer. The Jev values are the worked examples in Typesafe’s docs, one per primitive.

LLM
Jev

noul

Is the customer asking for a human agent?

LLM

“Yes, clearly.”

parse it to a boolean

Jev

noul 0.99

P(true), one number

0 0.5 1

choice

Which department: returns, shipping or billing?

LLM

“Returns, I think”

match it to an option name

Jev

choice: returns

confidence 1.0

1 0.5 0 shipping returns billing

score

How severe is the issue, on levels 0 to 2?

LLM

“Probably a 2”

a number as text, one value

Jev

score 1.43

confidence 0.35

1 0.5 0 1.43 0 1 2 level
Figure 2: An LLM answers each question with a string that code must then parse into a value. Jev returns a probability for every declared answer, shown here with the worked examples from Typesafe's docs.

Giving up string generation entirely is what buys the speed, the low price and the type guarantees. TypeSafe charges $0.042 per million input tokens, with output free, and reports response times of 70ms-500ms. On accuracy, its own benchmark is modest: Jev roughly ties GPT-5.6 Terra and trails GPT-5.6 Sol and Opus 5 by a few points. It trades a little accuracy for a lot of speed and cost. The bigger promise is calibration. The launch post's comparison table says "Calibrated: higher confidence means higher accuracy." The documentation defines confidence more narrowly: "a statistic computed from the probability distribution the answer already gives you," where 1.0 means all the mass sits on one option. One small independent test found that accuracy rose steadily with Jev's stated confidence, and its only mistake came back with the lowest confidence in the run.For securing agents, that fits the job nicely: because each answer returns in milliseconds with a probability attached, you can judge everything an agent sees, and when it's uncertain, the harness around Jev can escalate the decision.

Where it fits in security

In general, security seeks to answer one of four questions (see figure 3). Classify asks what something is, Decide asks what should happen, Plan asks what to do next over many steps, and Author asks what should exist: a rule, a policy, a report. Jev covers the first two. Planning needs parameters it can't produce, like a hostname discovered two steps earlier, and authoring needs text.

Four kinds of AI work in security

What a model is asked to produce, from a single label to a durable artifact.

Classify

What is this?

Output

Label or score

Examples

Phishing, malware, prompt injection

Decide

What should happen?

Output

One action

Examples

Allow or deny, quarantine, revoke

Plan

What next?

Output

Sequence of actions

Examples

Investigation, hunting, pentesting

Author

What should exist?

Output

Durable artifact

Examples

Detection rules, policies, reports

Figure 3: The questions asked in security settings.

Within those two, volume and depth decide the fit (see Figure 4). Decisions with a fixed set of answers, made millions of times a day, like policy checks, email verdicts and tool-call permits, are where Jev competes with rules and classic ML, because LLMs are priced out of that volume. Investigations and threat hunts happen a few times a day and usually involve writing new queries along the way, so an LLM wins there. The interesting zone lies in between: too many events for an LLM and too much ambiguity for a rule. Prompt injection screening and alert triage live there.

Where fast and cheap actually matters

Volumes are illustrative, order of magnitude.

Pinch to zoom in on the chart

Depth wins → LLMs Cost per call barely matters The gap → tier it Too many events for an LLM on each Anything works Often just rules or a human Fast & cheap wins → Rules, ML, Jev Latency and cost per call dominate Threat hunts Investigations Detection rule writing Alert triage Autonomous pentests Incident reports Prompt injection screening Containment decisions Vulnerability scoring Tool-call permits Email verdicts Malware verdicts Policy enforcement 1/day 10 100 1k 10k 100k 1M 10M 100M/day Decision volume Reasoning depth

Top right is the interesting zone: a cheap model screens, an LLM takes the residue.

Bottom right is where Jev competes, against rules and classic ML, not against LLMs.

Figure 4: Mapping security tasks

That's where the ML stack has shifted. A text classifier used to mean labeled data and a training run for every question. Now it's a question and a threshold. Jev replaces those classifiers, keyword rules and label-picking LLM calls, but not ML models like XGBoost on numeric features or an LLM that has to reason.

In practice, you compose tools rather than pick one (figure 5). Rules drop the known good and bad, Jev judges the rest, an LLM takes the ambiguous residue, and an analyst owns what reaches them.

Compose, don’t pick: a tiered pipeline

Each tier passes on only what it cannot settle. Volumes are illustrative.

Rules

10M events/day

Drop the known-good and known-bad

Microseconds, near free

Jev

1M events/day

Classify and decide on every remaining event

Fast and cheap, as claimed

LLM

10k events/day

Reason over the ambiguous residue, plan next steps

Seconds, real cost per call

Analyst

100 cases/day

Review, approve, own the outcome

Minutes to hours

Figure 5: Where Jev fits in the tiers of use.

What others have built

The projects built so far put Jev in the tier Figure 5 predicts: most run fixed rules first, ask Jev a few typed questions about whatever those rules let through, and send the uncertain cases to a person.

Jev Agent Action Guardrails

Filled dot means the tool has the trait. Dense columns are what they share; sparse columns are where they differ.

Has the trait Does not

Pinch to zoom in on the table

ToolJudgesRules liveJev primitiveDeployOutcomes
Tool call, before it runsTool result or contentProgress or done claimsWritten workIn the request payloadAs the question setFixed built-in questionsNoul (yes/no)Choice (pick one)Score (rubric)Rules run before the modelShadow mode by defaultLabels each gate can reach, including routing and ranking
abidenononoyesnoyesnoyesyesyesnonorepair / note / nothing
agent-chaperoneyesyesnonoyesnoyesyesyesyesyesyesallow / hold; results: pass / annotate / quarantine / redact
agi-jev-containmentyesyesyesnoyesnoyesyesyesnoyesnoallow / hold / refuse; incident level L0-L5
approval-judge-bridgeyesnononoyesnononoyesnononoapprove / escalate / deny
bounceryesnonoyesnoyesyesyesnonoyesyesallow / ask / deny
dsh-jev-interceptoryesnononononoyesyesyesyesyesyesdelegate / ask / deny; ranks recalled messages
Edward (Jev optional)nonoyesnononoyesnoyesnoyesnocontinue / pause / cancel / escalate / ask for approval
foremannonoyesnonoyesnoyesnonononosteer / verify / escalate / done
fxyesnononoyesnononoyesnononoclear / caution
jev-agent-authorizationyesnononononoyesyesyesyesyesnoallow / step up / deny
jev-axiyesyesyesyesnonoyesyesyesyesyesnoallow / ask / deny
jev-engineeringyesnonoyesnonoyesyesyesyesyesyesallow / ask / deny; routes PR review and model tier
jev-guardyesyesnonononoyesyesyesyesyesnoallow / ask / block
jev-lintyesnonoyesnoyesyesyesnonoyesnolikely / possible violation; block or hint
jev-playwright-mcpyesyesnonononoyesyesyesnoyesnoallow / blocked / needs confirmation
jev-reviewnononoyesnonoyesyesyesyesnonoranks severity 0-3; routes to a reviewer
jevguard-mcpyesnonoyesnonoyesyesyesyesnonoallow / ask / deny; patches: approve / revise / reject
pi-subagent-jevyesnonononoyesyesyesnonononoallow / block
pi-wardenyesyesyesyesyesyesnoyesyesyesyesnoallow / warn / confirm / steer / hold
r2r-jevyesnononononoyesyesnonononoallow / deny, after accept / hold / reject
reflexyesyesyesnononoyesyesyesyesyesnoallow / ask / block / steer; routes model tier
Ring Zero Securityyesyesyesnoyesnoyesyesyesnoyesyesallow / deny; Jev can only raise severity
SemanticPolicy (Jev optional)yesyesnononoyesnoyesyesyesnonoallow / warn / abstain / escalate / deny
snifftestnononoyesnoyesnoyesnonoyesnoflag / no judgment / pass
stingraynonoyesnononoyesyesnonononoblock / pass
toolgateyesnononoyesyesyesyesnonoyesnoallow / ask / deny / pass through
Figure 6: Jev-based projects so far that analyze individual agent actions.

Figure 6 compares open-source guardrail tools built on Typesafe's Jev model. You can find more projects on awesome-jev repos like https://github.com/yibie/awesome-jev. Every tool follows the same basic design: it asks Jev a small set of typed questions in one request and turns the answers into an action with a fixed threshold. Most of them judge a tool call before it runs, and most carry their rules as fixed questions built into the tool. The clearest differences are where install-specific rules live and how many outcomes a gate can reach. approval-judge-bridge and fx pass policy text alongside the input, abide and foreman turn the rules into the questions themselves, and the outcomes run from two labels in fx to five in pi-warden.

Jev Agent Multi-step Behavior Guardrails

Filled dot means the monitor has the trait. Tools marked (session checks) also gate single actions and appear in the action-gate figure.

Has the trait Does not

Pinch to zoom in on the table

MonitorUnit judgedData sourceTimingWatches forKeepsOutput
One eventChain of eventsSessionAcross sessionsAgent hooksKernel or OSWraps the agentLiveEnd of turnAfter the factSecurityReli­abilityWork verifi­cationGover­nanceEvent logGraphCounters or trustWhat the monitor can do or report
agi-jev-containmentyesyesyesyesnonoyesyesnoyesyesnononoyesyesyesincident level L0-L5; allow / hold / refuse
Cannyyesnoyesnoyesnonoyesyesnoyesnoyesnoyesnonorefuses "done" without a passing check; rule notes
Edward (Jev optional)nonoyesnononoyesyesnonoyesyesnonoyesnoyescontinue / pause / cancel / escalate
foremannonoyesnoyesnoyesyesnononoyesyesnoyesnonosteer / verify / escalate / done
jev-axi (supervise)nonoyesnoyesnonoyesyesnonoyesyesnoyesnonoconcern / done / not done
pi-warden (session checks)yesnoyesnoyesnonoyesyesnoyesyesyesnononoyessteer / hold; flags unbacked "done"
r2r-jevnonoyesyesnonoyesyesnononononoyesyesnoyeslowers trust, delegation, authorization
reflex (session checks)nonoyesnoyesnoyesnoyesnonoyesyesnononoyesnudge to continue / ask the user
Ring Zero Securityyesyesyesyesyesyesnoyesnoyesyesnononoyesnoyesblocks protected files; findings to SIEM
stingraynonoyesnoyesnononoyesnononoyesnoyesnonoblock with a nudge / pass
Figure 7: Jev-based projects analyzing multi-step agent trajectories.

Jev-based projects have also been set up to guard multi-step agent behavior, not just the single actions that were the subject of the previous figure. Only one of them was tested on public benchmarks: Edward ran StepShield, RedCode-Exec and ATBench, and even there Jev is an optional backend, while the others were tested on their own cases or not at all. All of them watch a single agent on a single machine, by hooking into it, wrapping it, or reading kernel events on its host, and none collects telemetry from a fleet of agents or analyzes it afterward across machines. Only two, agi-jev-containment and Ring Zero Security, link events into attack chains and track them across sessions. Both are days-old projects, one a hackathon demo, and the rest mostly check whether work is finished or stuck rather than whether it is risky.

Putting it to the test

Beyond Edward’s result, nobody has measured a Jev-based monitor on public data. We adapted part of Manifold’s system for analyzing agent behavior to use Jev as a model and ran in on two evaluations. The first is StepShield, which tests whether a monitor detects when an agent breaks the user’s stated contraints (Table 1). Our setup exceeds the performance of both SetpShield’s LLMJudge and Edward.

Detector EIR (k=3) Recall FPR
Manifold Jev 0.92 97.2% 5.6%
Edward + Jev 1.13 0.91 59.3% 10.2%
LLMJudge (GPT-4.1-mini) 0.89 95.4% 5.6%
HybridGuard 0.40 75.9% 44.4%
StaticGuard (847 rules) 0.23 86.1% 77.8%
ConstraintGuard 0.08 24.1% 25.0%
Random baseline 0.12 50.0% 50.0%
Table 1: StepShield evaluations. All rows are on test_holdout: 108 rogue and 108 clean trajectories. EIR (k=3), the early intervention rate, is the share of detected rogue trajectories whose first alert fires at the step where the agent goes rogue or within three steps after it; an alert before that step earns no credit. FPR is the share of clean trajectories flagged. The last five rows are from StepShield. The Manifold row runs the frozen setup step by step, as StepShield's own monitors do, and the metrics come from StepShield's own scoring code. Manifold’s Jev thresholds were calibrated on the training set. The Edward row is Edward's self-reported result.

We also tested this setup against the AgentDrift evaluation (Table 2). This eval tests if a monitor can determine whether an indirect prompt injection hijacked an agent. The authors of the AgentDrift eval trained a transformer-based classifier from scratch on their training set (DriftNet). The performance of this model shows the power of this type of model. Jev, on the other hand, shows its merit in its ability as a general model to approach the performance of the dedicated model, without any fine-tuning. Jev still has trouble with the ‘hard negative’ class, which are the benign cases that look like an attack.

Detector Recall FPR Benign Failed attack Hard negative
Manifold Jev, max F1 93.7% 5.8% 0.2% 1.8% 24.5%
Manifold Jev, 1% FPR 77.9% 0.5% 0.0% 0.0% 2.5%
DriftNet 98.5% 1.5% 1.5% 0.0% 2.9%
Logistic Regression (surface baseline) 57.9% 10.7% 9.0% 17.0% 8.3%
Table 2: AgentDrift evaluations. AgentDrift counts only successful attacks as positive. FPR is over all non-attack trajectories, and the last three columns are the flag rate on each non-attack class; a failed attack carries a real injection the agent resisted, and a hard negative carries legitimate content that looks like an attack. The max F1 and 1% FPR thresholds were chosen on the AgentDrift validation set. DriftNet is a transformer based classifier trained on AgentDrift's own training split, and its row and the surface baseline come from the DriftNet paper.

What Others Found

Our tests are one team’s results. Many others have been published. Here we go through two of the security-relevant cases: Vega ran Jev as a gate in front of its alert-triage agent, and Sardine replayed a fraud ring through a Jev pipeline.

What two early users found

Jev as a gate in front of a triage agent, and Jev against hand-written fraud rules. Numbers are the publishers’ own.

Alert triage (Vega): gate, don’t replace

Replayed alerts from two tenants; LLM-judged labels checked by analysts

Share of alerts Jev closed as a gate

Busy tenant
15%

98% of closures right

Quiet tenant
33%

99% of closures right

Accuracy when deciding every alert

Busy tenant
agent 72%
Jev 69%
Quiet tenant
agent 75%
Jev 65%

Jev figures use alert plus tenant context, its best setting.

Fraud (Sardine): rules-level recall without writing rules

One replayed fraud ring, 29 fraudulent transactions; no false-positive rate reported

Fraudulent transactions caught, of 29

Hand-written rules
26
Jev
27
LLM decides
18

The Jev pipeline cost 25 times less than the LLM pipeline.

Figure 8: Used as a gate, Jev closed 15 to 33% of Vega's alerts and was right 98 to 99% of the time, but it lost to the triage agent when it had to decide every alert. In Sardine's fraud replay, Jev caught 27 of 29 fraudulent transactions against 26 for hand-written rules, at a 25th of the LLM pipeline's cost.

Alert triage: gate, don't replace. Vega put Jev in front of its triage agent on replayed alerts from two tenants. Used as a one-sided gate, Jev closed 15% of alerts on the busy tenant and a third on the quiet one, and it was right about 98% of the time. Asked to decide every alert outright, it lost to the agent on both tenants. Its calibration was lopsided as well: on the busy tenant, a 0.9 on "escalate" was right only about six times in ten, while the confidence on its other answers held up. The results come from two tenants, with LLM-judged labels verified by analysts. The pattern to copy: clear what it’s confident about and escalate the rest.

Fraud detection: rules-level results without writing the rules. Sardine replayed a real fraud ring against a crypto on-ramp through three pipelines. Rules written by its analyst agent caught 26 of 29 fraudulent transactions, Jev caught 27, and an LLM making the call caught only 18. The Jev pipeline was also 25 times cheaper than the LLM one. The numbers Jev judged, like how many cards a user had used in the last day, were calculated in code first, which is exactly how TypeSafe says to use it. And with only 29 cases and no false-positive rate reported, Jev and the rules effectively tied. Jev matched hand-built rules without anyone having to build them.

Shortcomings

Jev can only create structured decisions and nothing else. Consequently, it’s not useful for generating any text: be it rules, reports or anything else which may need prose.

TypeSafe's own advice shows where Jev stops. Its guide says to keep control flow, fixed rules and actions in code, to break broad judgments into narrow questions, and to escalate uncertain cases to a person or a stronger model. TypeSafe also publishes a page listing where Jev is weak, which it calls model jaggedness. The model reads very literally, it's poor at math, counting and dates, and it gets worse with multi-step questions, irrelevant context or contradictory instructions (more details here: RedHub AI). Anything that isn't one of the predefined options, like a hostname found along the way, has to come from the surrounding code.

For security, two limits matter most. First, Jev's answer isn't a permission: if it says an action is safe, your system still has to check whether the action is allowed. Second, TypeSafe states plainly that Jev doesn't treat its input as hostile, and that an injected instruction or text arguing for its own classification can change the answer. In security, the input is almost always untrusted: emails, prompts, logs, web pages.

Jev alternatives

Jev launched on September 15th, and new alternatives have been appearing every day afterwards. They take four approaches. Some train a small open model to act like Jev, such as Kev on Qwen, openJev Verdict on ModernBERT, and AutoJev-27B on Qwen3.8. Some freeze an existing model and train only a small head on top, such as minojev and AnyJev. Some skip training and read the answer probabilities directly from an existing model, such as SemIf and Simple Jev. Others use a different kind of model underneath, most notably Laya, whose author says he published the same idea eighteen months before Jev. There is also a commercial rival, Nace.AI's Drex, which can be self-hosted.

The best alternatives now come close to Jev. In LangWatch's comparison of 11 tasks, Jev averaged 81.9%, against 80.0% for Eikos-27B, a finance-tuned Qwen. Shisa DE-1, built on Gemma 4 scored 79.7%; AutoJev-27B scored 78.9%. Two open models beat Jev on individual tasks: Eikos-27B on PII detection (95.2% against Jev’s 90.8%) and Shisa DE-1 on tool routing (81.3% against Jev’s 78.3%). LangWatch warns that many open models may have been trained on the test data. Nace reports Drex slightly ahead of Jev on the community Decision Index, but that is Nace's own run and not yet on the public board. TypeSafe's founder says the hard part isn't the model design but the training data that makes its confidence scores reliable (Firecrawl).

For security teams, the bigger question is where the model runs. Jev is only available as a service, from TypeSafe or through Cloudflare, so your emails, prompts and logs leave your environment. An open model can run on your own machines, where you can inspect it, retrain it on your own data, and test how it handles malicious input, though the strongest open contenders are 26B to 27B models that need a data-center GPU. Several open projects report well-calibrated confidence on their own tests, but none has shown it independently, and LangWatch found several open models whose confidence barely varied from one answer to the next. So the choice is simple: Jev still gives slightly better answers, and the open models and Drex give you more control.

Should you pay attention?

Yes, but for a narrower job than the demos suggest. Jev is not a smarter LLM. It's a fast, cheap way to make the same small judgment millions of times, with a confidence score that tells you when to double-check. In security, that describes much of the daily work.

Use Jev when Skip it when
The possible answers are known up front You need values it wasn’t offered, like a hostname found along the way
You make the same judgment at high volume The task needs math, counting or dates (use code)
The input is messy text: emails, prompts, alerts The signal is in numbers (use rules or classic ML, or compute them in code first)
You can escalate the uncertain cases The task needs several steps of reasoning or planning

The evidence so far points the same way. As a gate in front of a triage agent, Jev cleared a large share of alerts, but lost when it had to decide everything alone. In fraud detection, it matched hand-written rules without anyone writing them. The open-source tools built on Jev point the same way: nearly all of them use it as a gate in front of rules or a person, and only one has been tested on a public benchmark.

Two cautions before you build on it. Treat its answers as a signal, not a permission, and remember that TypeSafe itself says Jev can be steered by hostile input. In security, that's most input.

Our advice is to pick one high-volume decision you currently make with an LLM call, a regex or not at all, test Jev on your own data, and look closely at the confidence scores. If they reliably separate the easy cases from the hard ones, Jev will earn its place.

References

Typesafe

Explainers and independent tests

Architecture guesses

Benchmarks and public monitor results

What others found

Ecosystem index

  • yibie, 2026. "awesome-jev." The list the text links. Note: our references folder indexes cobanov/awesome-jev (21 tagged reconstructions); the two lists differ and the post currently points at yibie. GitHub repo

Variants

Commercial rivals and the LangWatch comparison

Hosted channels

Prior art and the other RLCD

About the author

  • Markus Gierlinger

    Markus Gierlinger

    Security Researcher

    Markus is a security researcher with over 12 years of experience across software engineering, security research, and product management.