AI decision models became a real product category in the space of 16 days. TypeSafe launched Jev on September 15, 2026, OpenAI previewed a Decisions API at DevDay on September 29, and on October 1 Cloudflare, Perplexity and Amazon all shipped their own, three of them with open weights.
They all do the same unusual thing: they never write a word. You give them a state and a list of typed questions, and they hand back probabilities.
If your agent spends tokens asking a frontier model “is this urgent, yes or no?”, a decision model answers the same question in tens to hundreds of milliseconds and returns a confidence you can actually threshold on.
I read the launch posts and docs from Cloudflare, Perplexity, AWS, OpenRouter’s Jev guide and the early independent comparisons. I have not benchmarked these models myself, so the numbers below are vendor-reported unless I say otherwise.
What Are AI Decision Models?
An AI decision model is a model trained to output a probability distribution over a fixed set of answers instead of generating text.
TypeSafe, which built Jev, calls these “System One” models, borrowing the fast, intuitive half of the fast-and-slow thinking split. A normal LLM is the slow half: it reasons in tokens and writes prose. A decision model is the snap judgement.
The input has two parts:
- State: the thing being judged, such as a support ticket, a log line, a pull request diff or a screenshot
- Questions: typed questions about that state, which the model answers in parallel
There are three question types, and the naming is shared across Jev and Clef:
| Type | What it returns | Example |
|---|---|---|
noul | Probability that a yes/no proposition is true | ”Is this request urgent?” |
choice | Probability for each option in a fixed list | ”Which team owns this: billing, technical or sales?” |
score | Probability-weighted rating on a scale | ”Rate this answer’s correctness from 1 to 5” |
(Source: Cloudflare Workers AI changelog and OpenRouter’s Jev guide, October 2026.)
Here is the shape of a call to Cloudflare Clef from a Worker, taken from Cloudflare’s changelog:
const response = await env.AI.run("@cf/cloudflare/clef", {
model: "clef",
state: "Checkout has been failing for every customer for the last hour.",
questions: {
urgent: {
type: "noul",
instructions: "Is this support request urgent?",
},
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: {
billing: "Payments, invoices, and refunds",
technical: "Outages, errors, and configuration",
sales: "Plans and upgrades",
},
},
},
});
The response contains a choice and a probability for every option. OpenRouter’s Jev guide shows the same structure: "choice": "billing", a probabilities map and a confidence value.
This is the part that I think matters most. With a text LLM, you parse a string, hope it matches one of your labels, and have no honest measure of certainty. TypeSafe claims Jev’s probabilities are calibrated, meaning that across many answers, a 0.8 is right about 80% of the time. If that holds, you can write if (p < 0.7) escalateToHuman() and mean it.
Decision models replace “parse the LLM’s answer and hope” with typed outputs and a confidence score your code can branch on.
Clef vs Jev vs Perplexity Decider vs Strands Decider
Five vendors are now in the category. Four have shipped something you can call today.
| Model | Vendor | Weights | Context | Input price per 1M |
|---|---|---|---|---|
| Jev 1.13 | TypeSafe | Proprietary | 32K | $0.042 |
| Clef | Cloudflare | Apache 2.0, 27B | 64K | $0.24 |
| Clef-flash | Cloudflare | Apache 2.0, 9B | 64K | $0.09 |
| pplx-decider-v1-27b | Perplexity | Apache 2.0, 27B | ~250K | $0.04 |
| Strands Decider 2B | AWS Strands Labs | Apache 2.0, ~1.9B | Not stated | Free, self-hosted |
(Sources: Cloudflare Workers AI docs and changelog; Perplexity Developers launch post; OpenRouter Jev guide; SiliconANGLE and MarkTechPost on Strands Decider, October 2026. Output tokens are free on all hosted options.)
TypeSafe Jev
Jev opened the category. It entered early access on September 15, 2026, and is on OpenRouter as typesafe/jev-1.13 and on Vercel AI Gateway. It is the only proprietary model in the table, with a 32K context window.
Jev also set the API shape. Cloudflare describes Clef as “fully Jev-API compatible”, which means the interface is turning into a de facto standard. That is good news if you want to switch vendors later.
Cloudflare Clef and Clef-flash
Clef is the speed play. Across 43 benchmark runs, Cloudflare reports a 38.8 ms median latency for Clef-flash and 209.3 ms for Clef, against 524.1 ms for Jev. Cloudflare says Clef leads Jev on 7 of 10 decision benchmarks, including 98.47 versus 95.75 on BFCL and 91.93 versus 88.19 on API-Bank.
Cloudflare’s own table also shows where Jev wins. On When2Call, Jev scored 80.97 and Clef 72.37. I give Cloudflare credit for publishing that row.
Clef adds vision. Per the Workers AI docs, you can send up to four images per request and ask between 1 and 64 questions per call. Cloudflare also announced an RL fine-tuning platform for training a decision model on your own data, but for now it runs as a hands-on engagement with its forward-deployed engineers, not a self-serve product.
One caution on latency. Nadir, a model-routing company, measured Clef-flash from the client side and got 191 to 205 ms, not 38.8 ms. Server-side numbers exclude network time. For a decision that gates a user request, the client-side figure is the one in your latency budget.
Perplexity pplx-decider-v1-27b
Perplexity is the price and context play. Its Decisions API costs $0.04 per million input tokens, the cheapest hosted option, and the model handles roughly 250K tokens of context because it is fine-tuned from Qwen3.8-27B. It accepts images too.
Perplexity reports 85.71% accuracy across an 11-benchmark panel of 7,210 samples, against 84.51% for Jev. Nadir’s comparison points out that pplx-decider wins the average by 1.2 points but loses 6 of the 11 individual rows. That gap between the headline number and the per-task numbers is the most useful thing to know about this whole category right now.
Strands Decider 2B
AWS went small. Strands Decider 2B starts from Qwen3.5-2B-Base with the language modelling head removed, so it physically cannot generate text. According to SiliconANGLE, it runs on a CPU, a consumer GPU or an Apple silicon Mac, and MarkTechPost reports decisions in about 115 ms. You install it with pip install strands-decider, which gives you a CLI and an HTTP server.
AWS says it ranks first on JevBench among public models that ship a full training recipe. That is a narrow claim, but the open recipe matters if you plan to fine-tune.
OpenAI Decisions API
OpenAI previewed its Decisions API at DevDay on September 29, built on a specialised version of GPT-6 Luna. In OpenAI’s demo it returned answers in about 150 ms, against 1.6 seconds for a standard Luna call. It is in limited preview with no published pricing, schema or docs yet, so I would not design around it this month.
For most teams, Perplexity Decider is the cheapest hosted default, Clef-flash is the fastest, and Strands Decider 2B is the one to self-host.
Decision Model Pricing vs a Frontier LLM
The cost story is more subtle than the launch posts suggest. Here is my arithmetic at 300 input tokens per decision, which is a short ticket plus two questions:
| Option | Cost per 1M decisions | Notes |
|---|---|---|
| pplx-decider-v1-27b | $12 | $0.04/M input, free output |
| TypeSafe Jev | $12.60 | $0.042/M input, free output |
| Clef-flash | $27 | $0.09/M input, free output |
| GPT-6 Luna as a text classifier | ~$35 | $0.10/$0.50, assuming ~10 output tokens |
| Clef | $72 | $0.24/M input, free output |
| Claude Sonnet 5.5 as a classifier | ~$700 | $2/$10, assuming ~10 output tokens |
(Pricing from vendor pages as of October 2026; per-decision math is my own estimate.)
Two things jumped out when I worked this through.
First, decision models are not dramatically cheaper than a small LLM. GPT-6 Luna already costs about the same as Clef-flash for short classification calls. The saving against Luna is latency and calibrated output, not money.
Second, the saving is enormous if you currently route decisions through your main agent model. Plenty of agent frameworks ask the same model that writes your code to also decide which tool to call. If that model is Claude Sonnet 5.5, moving the routing step to a decision model cuts its cost by more than 95% on my numbers. I compared the $2 mid-tier models in my GPT-6.1 Sol vs Claude Sonnet 5.5 breakdown, and both are wasted on yes/no questions.
Decision models pay off most when they replace a frontier model on routing, not when they replace an already-cheap small model.
Where Decision Models Fit in an Agentic Coding Stack
Decision models belong at the forks in an agent loop, not in the work itself.
If you have read my guide to agentic coding, you know an agent loop is mostly the model deciding what to do next. Many of those decisions have a fixed answer set:
- Model routing: does this task need the expensive model or the cheap one?
- Tool selection: should the agent search, read a file or run tests next?
- Guardrails: does this shell command touch production, yes or no?
- Review triage: is this diff low, medium or high risk?
- Evals: score this output from 1 to 5 against a rubric
That guardrail case is the one I find most interesting. A decision model with a 39 to 200 ms response time is fast enough to sit in front of every tool call without making the agent feel sluggish. The same idea applies to desktop agents like GitHub Copilot’s new computer use, where a quick yes/no check before each click is cheap insurance. In Claude Code, a Claude Code mod hooked on tool calls could call one to block risky commands. I have not built that, so treat it as a design idea rather than a recipe.
What decision models cannot do is anything open-ended. They will not write code, explain a stack trace or summarise a pull request. Every answer has to be on your list before you ask.
There is also a calibration caveat. Calibration is a property measured across many answers on a known distribution. On your own data, nobody has checked it yet, so log the probabilities next to outcomes for a few weeks before you trust a threshold.
Use a decision model wherever your code currently parses an LLM’s one-word answer, and keep the frontier model for the work that needs prose or code.
Which Decision Model Should You Use?
Here is my recommendation, based on what the vendors have published so far:
- Default hosted choice: Perplexity pplx-decider-v1-27b. It is the cheapest hosted option at $0.04 per million input tokens, has the longest context and handles images.
- Latency-critical paths: Cloudflare Clef-flash. It has the lowest published median latency, and it is the obvious pick if you already run on Workers.
- Self-hosted or air-gapped: Strands Decider 2B. It is small enough for a laptop and ships with a full training recipe.
- Jev if you want the original API and its calibration claims, and you do not need open weights.
- Wait on OpenAI’s Decisions API until pricing and docs are public.
Because Clef is Jev-API compatible and the others follow the same state-and-questions shape, switching costs are low. Write a thin wrapper, run your own labelled sample through two of them, and pick on your data rather than on anyone’s 11-benchmark average.
AI decision models are the first 2026 model category that makes agents cheaper and faster at the same time, as long as you only ask them questions with a fixed answer.