
Jev, system one models, and how to use them in security
Cody NashResearcher
Markus GierlingerSecurity Researcher
TL;DR
- What it is: Jev isn't a smarter LLM. It answers narrow, predefined questions in milliseconds, for almost nothing, and says how sure it is. It has launched a new wave of innovation in the fast, low-cost model space.
- Where it fits in security: use it as a filter. It clears the cases it’s confident about; everything else goes to rules, a person, or a stronger model.
- Before you build on it: attackers can steer it through its input, and it only runs as a hosted service. Test it on your own data first, or try an open alternative, which now scores within a few points of Jev.
In the beginning
Jev was the first of a new wave of “system one” models: they think fast and cost little. They are an alternative to the large language models that run their entire process on every token. LLMs are the heavy hammer that has been applied to every problem. But for applications with a high volume of tasks, including checking for prompt injection or monitoring logs of agent behavior, LLMs can be too expensive. Expert systems and classical machine learning covered the easy share of these tasks. This new wave of system one models leverage the generalizability of the transformer architecture to bring more intelligence to the rest.
Security runs on small decisions made at high volume
Is this email phishing? Is this prompt an injection attempt? Should this agent be allowed to make this tool call? Does this alert deserve an analyst? None of these are hard, but all of them happen millions of times.
Until now, those decisions lived in rules and workflows (i.e. Security Orchestration, Automation, and Response (SOAR) playbooks), which break on anything unanticipated; in trained classifiers, which need labeled data for every question; or in LLMs, which are too slow and expensive to run on every event. Jev claims a fourth option: a model's judgment at a classifier's speed and cost. It handles the messy, ambiguous text that breaks a rule, like a phishing email with no known indicators or a prompt phrased to slip past a regex. The workflow still decides what to ask and what to do with the answer. Jev takes over the brittle decision logic inside it; the playbook stays.
Behind the curtain (what Jev is)
TypeSafe calls Jev its first System One model: it's built for fast, intuitive decisions inside software, not for slow, deliberate reasoning. It doesn't write text. You give it a block of input (the state) and a predefined set of questions, and it answers them all in one parallel call instead of one token at a time. Each answer is a Choice among up to 255 options, a Score on a scale you define, or a Noul, a yes-or-no question returned as a probability (more details here: thoughts.jock.pl, firecrawl).
Giving up string generation entirely is what buys the speed, the low price and the type guarantees. TypeSafe charges $0.042 per million input tokens, with output free, and reports response times of 70ms-500ms. On accuracy, its own benchmark is modest: Jev roughly ties GPT-5.6 Terra and trails GPT-5.6 Sol and Opus 5 by a few points. It trades a little accuracy for a lot of speed and cost. The bigger promise is calibration. The launch post's comparison table says "Calibrated: higher confidence means higher accuracy." The documentation defines confidence more narrowly: "a statistic computed from the probability distribution the answer already gives you," where 1.0 means all the mass sits on one option. One small independent test found that accuracy rose steadily with Jev's stated confidence, and its only mistake came back with the lowest confidence in the run.For securing agents, that fits the job nicely: because each answer returns in milliseconds with a probability attached, you can judge everything an agent sees, and when it's uncertain, the harness around Jev can escalate the decision.
Where it fits in security
In general, security seeks to answer one of four questions (see figure 3). Classify asks what something is, Decide asks what should happen, Plan asks what to do next over many steps, and Author asks what should exist: a rule, a policy, a report. Jev covers the first two. Planning needs parameters it can't produce, like a hostname discovered two steps earlier, and authoring needs text.
Within those two, volume and depth decide the fit (see Figure 4). Decisions with a fixed set of answers, made millions of times a day, like policy checks, email verdicts and tool-call permits, are where Jev competes with rules and classic ML, because LLMs are priced out of that volume. Investigations and threat hunts happen a few times a day and usually involve writing new queries along the way, so an LLM wins there. The interesting zone lies in between: too many events for an LLM and too much ambiguity for a rule. Prompt injection screening and alert triage live there.
That's where the ML stack has shifted. A text classifier used to mean labeled data and a training run for every question. Now it's a question and a threshold. Jev replaces those classifiers, keyword rules and label-picking LLM calls, but not ML models like XGBoost on numeric features or an LLM that has to reason.
In practice, you compose tools rather than pick one (figure 5). Rules drop the known good and bad, Jev judges the rest, an LLM takes the ambiguous residue, and an analyst owns what reaches them.
What others have built
The projects built so far put Jev in the tier Figure 5 predicts: most run fixed rules first, ask Jev a few typed questions about whatever those rules let through, and send the uncertain cases to a person.
Figure 6 compares open-source guardrail tools built on Typesafe's Jev model. You can find more projects on awesome-jev repos like https://github.com/yibie/awesome-jev. Every tool follows the same basic design: it asks Jev a small set of typed questions in one request and turns the answers into an action with a fixed threshold. Most of them judge a tool call before it runs, and most carry their rules as fixed questions built into the tool. The clearest differences are where install-specific rules live and how many outcomes a gate can reach. approval-judge-bridge and fx pass policy text alongside the input, abide and foreman turn the rules into the questions themselves, and the outcomes run from two labels in fx to five in pi-warden.
Jev-based projects have also been set up to guard multi-step agent behavior, not just the single actions that were the subject of the previous figure. Only one of them was tested on public benchmarks: Edward ran StepShield, RedCode-Exec and ATBench, and even there Jev is an optional backend, while the others were tested on their own cases or not at all. All of them watch a single agent on a single machine, by hooking into it, wrapping it, or reading kernel events on its host, and none collects telemetry from a fleet of agents or analyzes it afterward across machines. Only two, agi-jev-containment and Ring Zero Security, link events into attack chains and track them across sessions. Both are days-old projects, one a hackathon demo, and the rest mostly check whether work is finished or stuck rather than whether it is risky.
Putting it to the test
Beyond Edward’s result, nobody has measured a Jev-based monitor on public data. We adapted part of Manifold’s system for analyzing agent behavior to use Jev as a model and ran in on two evaluations. The first is StepShield, which tests whether a monitor detects when an agent breaks the user’s stated contraints (Table 1). Our setup exceeds the performance of both SetpShield’s LLMJudge and Edward.
We also tested this setup against the AgentDrift evaluation (Table 2). This eval tests if a monitor can determine whether an indirect prompt injection hijacked an agent. The authors of the AgentDrift eval trained a transformer-based classifier from scratch on their training set (DriftNet). The performance of this model shows the power of this type of model. Jev, on the other hand, shows its merit in its ability as a general model to approach the performance of the dedicated model, without any fine-tuning. Jev still has trouble with the ‘hard negative’ class, which are the benign cases that look like an attack.
What Others Found
Our tests are one team’s results. Many others have been published. Here we go through two of the security-relevant cases: Vega ran Jev as a gate in front of its alert-triage agent, and Sardine replayed a fraud ring through a Jev pipeline.
Alert triage: gate, don't replace. Vega put Jev in front of its triage agent on replayed alerts from two tenants. Used as a one-sided gate, Jev closed 15% of alerts on the busy tenant and a third on the quiet one, and it was right about 98% of the time. Asked to decide every alert outright, it lost to the agent on both tenants. Its calibration was lopsided as well: on the busy tenant, a 0.9 on "escalate" was right only about six times in ten, while the confidence on its other answers held up. The results come from two tenants, with LLM-judged labels verified by analysts. The pattern to copy: clear what it’s confident about and escalate the rest.
Fraud detection: rules-level results without writing the rules. Sardine replayed a real fraud ring against a crypto on-ramp through three pipelines. Rules written by its analyst agent caught 26 of 29 fraudulent transactions, Jev caught 27, and an LLM making the call caught only 18. The Jev pipeline was also 25 times cheaper than the LLM one. The numbers Jev judged, like how many cards a user had used in the last day, were calculated in code first, which is exactly how TypeSafe says to use it. And with only 29 cases and no false-positive rate reported, Jev and the rules effectively tied. Jev matched hand-built rules without anyone having to build them.
Shortcomings
Jev can only create structured decisions and nothing else. Consequently, it’s not useful for generating any text: be it rules, reports or anything else which may need prose.
TypeSafe's own advice shows where Jev stops. Its guide says to keep control flow, fixed rules and actions in code, to break broad judgments into narrow questions, and to escalate uncertain cases to a person or a stronger model. TypeSafe also publishes a page listing where Jev is weak, which it calls model jaggedness. The model reads very literally, it's poor at math, counting and dates, and it gets worse with multi-step questions, irrelevant context or contradictory instructions (more details here: RedHub AI). Anything that isn't one of the predefined options, like a hostname found along the way, has to come from the surrounding code.
For security, two limits matter most. First, Jev's answer isn't a permission: if it says an action is safe, your system still has to check whether the action is allowed. Second, TypeSafe states plainly that Jev doesn't treat its input as hostile, and that an injected instruction or text arguing for its own classification can change the answer. In security, the input is almost always untrusted: emails, prompts, logs, web pages.
Jev alternatives
Jev launched on September 15th, and new alternatives have been appearing every day afterwards. They take four approaches. Some train a small open model to act like Jev, such as Kev on Qwen, openJev Verdict on ModernBERT, and AutoJev-27B on Qwen3.8. Some freeze an existing model and train only a small head on top, such as minojev and AnyJev. Some skip training and read the answer probabilities directly from an existing model, such as SemIf and Simple Jev. Others use a different kind of model underneath, most notably Laya, whose author says he published the same idea eighteen months before Jev. There is also a commercial rival, Nace.AI's Drex, which can be self-hosted.
The best alternatives now come close to Jev. In LangWatch's comparison of 11 tasks, Jev averaged 81.9%, against 80.0% for Eikos-27B, a finance-tuned Qwen. Shisa DE-1, built on Gemma 4 scored 79.7%; AutoJev-27B scored 78.9%. Two open models beat Jev on individual tasks: Eikos-27B on PII detection (95.2% against Jev’s 90.8%) and Shisa DE-1 on tool routing (81.3% against Jev’s 78.3%). LangWatch warns that many open models may have been trained on the test data. Nace reports Drex slightly ahead of Jev on the community Decision Index, but that is Nace's own run and not yet on the public board. TypeSafe's founder says the hard part isn't the model design but the training data that makes its confidence scores reliable (Firecrawl).
For security teams, the bigger question is where the model runs. Jev is only available as a service, from TypeSafe or through Cloudflare, so your emails, prompts and logs leave your environment. An open model can run on your own machines, where you can inspect it, retrain it on your own data, and test how it handles malicious input, though the strongest open contenders are 26B to 27B models that need a data-center GPU. Several open projects report well-calibrated confidence on their own tests, but none has shown it independently, and LangWatch found several open models whose confidence barely varied from one answer to the next. So the choice is simple: Jev still gives slightly better answers, and the open models and Drex give you more control.
Should you pay attention?
Yes, but for a narrower job than the demos suggest. Jev is not a smarter LLM. It's a fast, cheap way to make the same small judgment millions of times, with a confidence score that tells you when to double-check. In security, that describes much of the daily work.
The evidence so far points the same way. As a gate in front of a triage agent, Jev cleared a large share of alerts, but lost when it had to decide everything alone. In fraud detection, it matched hand-written rules without anyone writing them. The open-source tools built on Jev point the same way: nearly all of them use it as a gate in front of rules or a person, and only one has been tested on a public benchmark.
Two cautions before you build on it. Treat its answers as a signal, not a permission, and remember that TypeSafe itself says Jev can be steered by hostile input. In security, that's most input.
Our advice is to pick one high-volume decision you currently make with an LLM call, a regex or not at all, test Jev on your own data, and look closely at the confidence scores. If they reliably separate the easy cases from the hard ones, Jev will earn its place.
References
Typesafe
- Typesafe AI, 2026. "Introducing System One models and Jev." typesafe.ai
- Typesafe AI, 2026. "Typesafe documentation." docs.typesafe.ai
- Typesafe AI, 2026. "Confidence." Defines confidence as a statistic of the probability distribution over the answer, with no accuracy claim. docs.typesafe.ai/confidence
- Typesafe AI, 2026. "Model jaggedness: Jev 1.13." docs.typesafe.ai
- Typesafe AI, 2026. "How to build with System One." docs.typesafe.ai
Explainers and independent tests
- Build Fast with AI, 2026. "Jev AI review." blog.buildfastwithai.com
- jock.pl, 2026. "Jev (Typesafe) System One model benchmark." The small independent test the post cites twice. thoughts.jock.pl
- DataCamp, 2026. "System one models: Jev." datacamp.com
- Firecrawl, 2026. "What is Jev." Source of the founder's training-data remark quoted in Jev alternatives. firecrawl.dev
Architecture guesses
- Rogge, 2026. "Curious how the viral Jev model by Typesafe works." LinkedIn
- Gulli, 2026. "LinkedIn post." Open-source reconstructions of Jev, naming SemIf and jevlike. LinkedIn
Benchmarks and public monitor results
- glo26, 2026. "StepShield." Benchmark of agents breaking user constraints. GitHub repo
- Asif-0209, 2026. "AgentDrift." Indirect prompt-injection hijack eval. GitHub repo
- Pinjari and Saint-Germain, 2026. "DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents." arXiv:2609.10892
- VeridicalTech, 2026. "Edward." The project whose self-reported Jev 1.13 StepShield result Table 1 reproduces. GitHub repo
What others found
- Vega, 2026. "Jev closed up to a third of our triage agent's alerts." vega.io
- Sardine, 2026. "Fraud Eval on Jev." LinkedIn
Ecosystem index
- yibie, 2026. "awesome-jev." The list the text links. Note: our references folder indexes cobanov/awesome-jev (21 tagged reconstructions); the two lists differ and the post currently points at yibie. GitHub repo
Variants
- Nandakishor M (Convai Innovations), 2026. "laya." Priority claim: "I Built Non-Autoregressive Decision Models with RL a Year Ago." GitHub repo
- Palmer, 2026. "kev." Coverage: The New Stack, "Kev skips text generation." GitHub repo
- heman10x, 2026. "openJev-verdict-2.0." Model: rlcd-modernbert-151m. GitHub repo
- denis-pplx, 2026. "AutoJev." AutoJev-27B, LangWatch's 78.9% open model. GitHub repo
- zeredy879, 2026. "minojev." GitHub repo
- MorrisZJ, 2026. "AnyJev." GitHub repo
- TheoLeeCJ, 2026. "SemIf." GitHub repo
- Featherless AI, 2026. "simple-jev." GitHub repo
- Gundala, 2026. "Qwen-2.5-1B-RLCD." The Hugging Face weights the architecture reconstructions reverse-engineered. huggingface.co
- Pujari, 2026. "Dev." Bidirectional cross-encoder with candidate choices as prompt text. LinkedIn
Commercial rivals and the LangWatch comparison
- LangWatch, 2026. "Jev vs all." The 11-task comparison; contamination caveat applies. langwatch.ai
- Vicentino, 2026. "Eikos-27B." Finance-tuned Qwen, 80.0% average on the LangWatch board. huggingface.co
- Shisa AI, 2026. "Shisa DE-1." Built on Gemma 4, 79.7% average. huggingface.co
- Nace.AI, 2026. "Drex." Commercial self-hostable rival; the Decision Index claim is Nace's own run. nace.ai
- webdevtodayjason, 2026. "Decision Index." Community benchmark. GitHub repo
Hosted channels
- Cloudflare, 2026. "Typesafe Jev." The second hosted channel. developers.cloudflare.com
Prior art and the other RLCD
- Yang, Klein, Celikyilmaz, Peng and Tian, 2023. "RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment." ICLR 2024. The acronym collision only; unrelated to decision models. arXiv:2307.12950
- Nandakishor M et al., 2025a. "SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization." The March 2025 paper Laya's priority claim rests on. arXiv:2503.23303
- Nandakishor M et al., 2025b. "Confidence-Aware Routing for Large Language Model Reliability Enhancement." The second paper Laya cites as formalizing its framework. arXiv:2510.01237
- Hurn-Maloney, 2026. "GLiNER comparison post." LinkedIn
About the author

Security Researcher
Markus is a security researcher with over 12 years of experience across software engineering, security research, and product management.





