A decision model is an AI model that takes a state (text, JSON, and for some models images) plus a typed question with its allowed answers, and returns a probability for each answer in one pass. It generates no text. TypeSafe’s Jev and Cloudflare’s Clef are the two we at Startrise have measured, in our 168-decision Jev vs Clef comparison.
This page covers the AI sense, which arrived with TypeSafe’s Jev on September 15, 2026 and Cloudflare’s Clef on October 1, 2026. It is not DMN, Decision Model and Notation, the Object Management Group’s business-rules standard; not the decision model of decision theory; not a decision tree; and not a Markov decision process. Same words, four objects; a “decision model AI” search means this one.
Key Takeaways
- A decision model returns probabilities over answers you declare; an LLM writes text and can be asked to format it as JSON. Different objects, different failure modes
- In our 168-decision pilot, Jev 1.13 agreed with references on 149/167 (89.2%) and Clef on 150/167 (89.8%): one decision apart, inside noise
- Price per decision, metered: Jev $17 per million, Clef $63, Claude Opus 5.5 $2,708; a keyword rule, $0
- The probability is the product: Jev averaged 0.86 confidence when right and 0.69 when wrong (n=149/18); chat models prompted for JSON returned 1.00 either way
- Limits are real: on injection every decision model scored 4/12 or worse against a regex’s 12/12 (recall-only); TypeSafe’s own docs list counting, dates and option order as jagged edges
What is a decision model?
A decision model has three parts: a state, typed questions, and a distribution over the answers you allowed. TypeSafe’s documentation puts it in one line: “Jev evaluates typed questions against a state and returns structured results directly. No text generation, no parsing.”
The state is the evidence: a ticket, a post, a record, a page. Jev accepts text or JSON, up to 32k tokens of state inside a 64k-token request; Clef accepts text, JSON, images and video, up to 65,536 tokens.
The question is typed, and Clef and Liquid AI’s d1 share TypeSafe’s three primitives: a noul asks whether a statement about the state is true and returns a value from 0 to 1; a choice picks one option from a list you supply; a score places the state on an ordered rubric you define. Several questions ride in one request, “evaluated in parallel and in isolation” per TypeSafe’s docs; Clef takes up to 64.
The answer is a probability for each allowed value, plus a confidence number from 0 to 1 that summarises how concentrated the distribution is: 1.0 when all the probability sits on one outcome, 0 when it is spread evenly. Cloudflare’s launch post gives the shorter version: “A decision model makes classifications to help agents decide how to act, based on certain probabilities.” This is the layer underneath agents that Startrise AI Labs studies.
How is a decision model different from an LLM?
Decision model vs LLM comes down to the object that comes back. A chat model generates tokens. Ask it for a decision and it writes a document with the decision inside, and documents run out of room. A decision model returns the decision.
Our pilot’s failure log is the whole argument. Of 843 attempts, 3 failed, all on chat models: feed / AI tools / Opus 5.5 returned HTTP 402, a provider credit pre-check, retried and settled; feed / Developer tools / GPT-6 Luna spent 150 of a 250-token output cap on reasoning and left no visible decision; ads / Busy operator / Opus 5.5 truncated its structured JSON at the 250-token ceiling. Jev and Clef: zero errors, zero retries across 336 calls.
Then the uncertainty signal. A chat model’s “confidence” field is generated text; in the pilot GPT-6 Luna and Opus 5.5 returned 1.00 whether right or wrong (n=150/17 and 151/16). A decision model’s probability is its actual output, and Jev’s separated right from wrong by 0.17 on average. LLMs “tend to be overconfident” when verbalising confidence (Xiong and colleagues, ICLR 2024).
One number now, the rest below: Opus 5.5 cost 162x Jev per decision for 2 more correct decisions in 167. The task-by-task breakdown is our 168-decision Jev vs Clef comparison.
How is it different from a classifier, or from JSON mode?
A classic classifier is trained on labelled examples for a fixed label set; change the labels and you’re collecting data and retraining. Red Hat’s October 2026 benchmark had a pre-trained DeBERTa classifier ahead of Jev on prompt injection.
A decision model needs no training data. The label set is whatever question you send; to change it, edit the request.
JSON mode and structured outputs constrain an LLM to a schema. You get the shape, no usable probability, and still pay generation prices and latency. A schema guarantees well-formed text; it’s still text, and the 250-token truncation in our pilot happened inside a schema-constrained response.
| Property | Decision model | LLM with JSON mode | Classic classifier |
|---|---|---|---|
| Output | Probability per allowed answer plus confidence | Text shaped to a schema | Label plus score |
| Uncertainty signal | The probability itself | A “confidence” field that is generated text | A score that may need calibrating |
| Change the labels | Edit the question in the request | Edit the prompt or schema | Collect data and retrain |
| Training data you supply | None | None | Required |
| Writes prose, code or a rationale | No | Yes | No |
| Failure mode seen in our pilot | Wrong with low confidence (Jev 0.69 avg when wrong, n=18); confidently wrong on Clef Flash (0.85 avg when wrong, n=29) | Output-cap truncation (2 of 843 attempts) and a 402 pre-check | Not measured; the $0 keyword rule stood in: 12/12 on injection, 15/18 on feed |
| Metered cost per million decisions in our pilot | Jev $17, Clef $63 | GPT-6 Luna $27, Claude Opus 5.5 $2,708 | Not measured; keyword rule $0 |
| Best when | Bounded answer consumed by code; labels change often | The answer is text, or a person must read a reason | Labels are stable and you have thousands of labelled examples |
Pilot cells are agreement with post-hoc references on 6 to 12 record batches. Source: Startrise decision-model pilot, plan jev-pilot-2026-10-04.
With a stable label set and thousands of examples, the classifier may win. The pilot cells rest on post-hoc references from two blind AI reviewers; see why we distrust AI judges enough to use two blind reviewers.
How does a decision model work?
The model is post-trained to produce a distribution over the answers you declared instead of the next token. Cloudflare describes Clef as Qwen backbones “post-trained … to suit decision model use cases”; TypeSafe has a new training algorithm it hasn’t published.
The shape, trimmed to one question from TypeSafe’s quickstart. The request:
{
"state": "...",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
}
}
}
And the response:
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
}
}
}
The equivalent line on a chat model, with OpenAI’s Structured Outputs:
// Chat model: the shape of the text is guaranteed; the text is still generated
text: { format: { type: "json_schema", strict: true, schema } }
Then you threshold the probability and branch. Our pilot used a single 0.5 threshold, untuned. TypeSafe’s confidence page recommends tiers that rise with the cost of a mistake, “a confidence threshold is not one number”, and its example routes anything below 0.5 to a person: “Model is genuinely unsure. Don’t guess.”
Thresholds only work if the probabilities mean something. Calibration is the word for that: of all the answers given at 0.8, about 80% should be correct. TypeSafe says System One models “are trained for calibrated decisions”. A vendor claim, and TypeSafe itself adds the caveat: “Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.”
What we measured is the gap in average reported confidence. Jev 0.86 when right against 0.69 when wrong (n=149/18); Clef 0.94 against 0.84 (n=150/17); Clef Flash 0.91 against 0.85 (n=138/29); GPT-6 Luna and Claude Opus 5.5, 1.00 either way. A gap on a small pilot, not a calibration curve; we tuned nothing. Every record, probability and reference label is public.
What are Jev and Clef?
Jev 1.13 is TypeSafe AI’s hosted, closed-weight decision model, launched September 15, 2026; jev-1.13.0 is the current model per TypeSafe’s models page: $0.042 per million input tokens, output free, 64k tokens per request with 32k for state, text only. Not the Japanese encephalitis virus, not the YouTuber, the other “what is Jev” results: “We named Jev after William Stanley Jevons,” the economist, says TypeSafe’s launch post.
Clef and Clef Flash are Cloudflare’s open-weight decision models on Workers AI, released October 1, 2026 under Apache 2.0: Clef 27B on a Qwen3.8-27B backbone, Clef Flash 9B on Qwen3.5-9B. The weights are public; the training pipeline isn’t, or, as the Hacker News thread put it, “open weights, not open source”. Both take a 65,536-token context, text, JSON, images and video, up to 64 questions per request, at $0.24 and $0.09 per million input tokens, no output charge. Cloudflare calls them “fully Jev-API compatible”; the changelog says you switch by changing the endpoint and model.
Compatible doesn’t mean interchangeable. One explainer’s FAQ says the two “take the same request and return the same answers”. We sent identical requests to both and they disagreed on 12 of 168 decisions: Clef took links 9/9 to Jev’s 5/9, Jev took ads 15/18 to Clef’s 13/18; see where Jev and Clef disagreed, record by record.
TypeSafe claims “193.6x faster, 444.6x cheaper” than the average of GPT-6 Astra and Fable 5.1, measured by TypeSafe on TypeSafe’s evals. Cloudflare prints medians of 209.3 ms for Clef, 38.8 ms for Clef Flash and 524.1 ms for Jev, measured by Cloudflare on Cloudflare’s infrastructure. Ours, end to end through our gateway, one timing trial: Jev 320 ms serialized per decision, Clef 655 ms, Clef Flash 316 ms. Different distances; both can be true.
The field is wider than two names: Vercel’s explainer lists seven more, from Perplexity’s pplx-decider-v1-27b to OpenAI’s Decisions API (announced September 29, 2026, limited preview), and Red Hat has run DiffusionGemma on vLLM. We measured none of these; the Clef release is logged in our model tracker’s Clef release entry.
What is a System One model?
“System One model” is TypeSafe’s category name. The concept page defines it as “a class of AI models built to make fast, structured decisions that software can use directly” and says what they don’t do: “do not write replies, produce code, or generate explanations of their reasoning.” The Kahneman borrowing, fast System 1 against slow System 2, is TypeSafe’s own framing.
The phrase has also become an API name. Cloudflare’s changelog says “Clef follows the System One API”; Liquid AI’s decision endpoint is literally /decisions/v1/systemone. In practice “decision model” is the category; “System One” is TypeSafe’s brand for it and the request shape others copied.
What can you use a decision model for?
Anything that ends in a branch: routing requests to agents, queues or humans; guardrails for out-of-scope, abusive or injected text; scoring against a rubric; screening a batch so only the shortlist reaches an expensive LLM; deciding whether a loop should stop.
In the pilot, filter a feed: Jev and Clef 17/18. Label a stream: launches, Jev 18/18; bursts, 25/26 for both. Rank a shortlist: rentals 36/36 on every row; ads, Jev 15/18 and Clef 13/18; links, Clef 9/9 and Jev 5/9. Check grounding against a source: 12/12 for both on RAGTruth human labels. Flag leaks and rule breaches: leaks 6/6 on every row; rules 11/12 for both.
The pattern that matters most is the cascade, or two-tier guardrail. Run the cheap decision model on everything. Above the high threshold, act. Between thresholds, escalate to a frontier model. Below the low threshold, hand to a person.
The top tier is what we add to live agents as AI Monitoring & Guardrails (from $2,500); the bottom tier is covered in approval workflows for agents and human-in-the-loop over MCP; the feed and launches class is what AI Content Pipelines ($5,000) builds.
What are the limits of decision models?
The vendor’s list. TypeSafe’s own jaggedness page for Jev 1.13, reviewed October 2, 2026, is candid for a vendor: “jev-1.13 does not count reliably”; “jev-1.13 reads dates as text, not as ordered quantities”; “Accuracy falls as the state grows with content unrelated to the decision”; “the order of a Choice’s options can affect the answer, and jev-1.13 leans toward the option that comes first”; adversarial content “can move the answer” because the model “does not treat it as hostile by default”; “Scoping words, negations, and implied conditions are read at face value.”
The independent benchmark. On October 2, 2026, Red Hat Developers (Rob Geada, Mac Misiura, Shelton Cyril) ran Jev 1.13 against eight guardrail models on prompt-injection and content-safety sets; Clef was not included. Their conclusion, verbatim: “decision models like Jev do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy.”
Our pilot. Injection was the category’s worst task: Jev and Clef 4/12, Clef Flash 3/12, Opus 5.5 8/12, the $0 keyword baseline 12/12. Recall-only: every record is an attack, so a filter that flags everything also scores 12/12. Confidently wrong exists: Clef Flash averaged 0.91 when right and 0.85 when wrong (n=138/29), too narrow for a threshold to catch. Both decision models made the identical false positive on the feed Developer-tools preset (“Remove and Disable Apple macOS 27 AI Models Tool”). And one limit raised on the Clef thread: no way to pass a prior or fine-tune through the hosted API (Cloudflare sells RL fine-tuning as a service; we did not test it). The pilot’s own sentence stands: “These small pilots do not establish general accuracy.”
How much does a decision model cost?
Vendor rate cards first: Jev $0.042 per million input tokens, Clef $0.24, Clef Flash $0.09, no output charge on any. Cheap chat models sit in the same input band, so token prices tell you little. Price the decision: output is free and the input is just the evidence.
Metered on 168 decisions each, the bill was Jev $0.0028 ($17 per million decisions), Clef Flash $0.0040 ($24), GPT-6 Luna $0.0045 ($27), Clef $0.0106 ($63) and Claude Opus 5.5 $0.4550 ($2,708). Clef is 3.8x Jev; Opus 5.5 is 162x Jev and 43x Clef.
The card says Clef should be 5.7x Jev ($0.24 against $0.042). The meter says 3.8x, because Clef counted fewer input tokens for the same evidence: 44,124 against Jev’s 66,605. The whole pilot, 843 attempts, cost $0.4769; Opus 5.5 was $0.4550 of it. Printing the bill next to the score is a habit from the Startrise LLM Benchmark.
When should you use an LLM instead?
When the output is prose, code, a summary or a plan: System One models “do not write replies, produce code, or generate explanations”. Also when a person has to read a written reason with every decision; a probability isn’t an explanation.
When the question needs several steps of reasoning, arithmetic or a date comparison; the jaggedness list says the model reads dates as text and doesn’t count reliably. When your rule already wins, as the regex did on injection at 12/12 for $0. And when your labels are stable and you have thousands of labelled examples: train a classifier, as Red Hat’s benchmark argues.
Most production systems run both. The LLM generates and plans; decision models sit at the branches and pick the tool, queue or human. That’s what we build as AI Agents That Act (from $3,500): agents on Mastra with routing, fallbacks, guardrails and approval gates, where the model at each branch ships with its evidence.
How do we know? The 168-decision pilot
The numbers come from one first-party pilot, plan jev-pilot-2026-10-04, completed October 4, 2026 at 23:06 UTC: 10 tasks, 26 presets, 130 batches, 840 completed calls, 3 failures, 843 attempts, $0.4769 spent. Six rows made 168 decisions each: Jev 1.13, Clef, Clef Flash, GPT-6 Luna and Claude Opus 5.5 prompted for JSON under a 250-token output cap, and a local keyword baseline with no API call. 167 decisions were reviewed per API model, 155 for the baseline, which has no grounding run.
Every row received identical evidence and question per preset, at most 1,400 serialized request bytes, threshold 0.5. References came after the fact from two blind AI reviewers who never saw model outputs, accepted only on consensus; rules and leaks were deterministic; grounding used RAGTruth’s human annotations; injection used BIPIA attack labels.
Caveats: agreement is not accuracy; batches are 6 to 12 records; one decision is 0.6 points; no single number is the score. The hands-on decision-model study shows the run as a recorded replay, not inference time; every record, decision, reference, cost and failure is on the evidence page; the full Jev vs Clef write-up has the task-by-task split.
How do you try a decision model?
Three routes. Jev through TypeSafe’s API, starting at the official quickstart. Clef and Clef Flash on Cloudflare Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash, with 10,000 free neurons per day, or self-hosted from the Apache 2.0 weights on Hugging Face. Or both through a gateway for one bill; OpenRouter lists Cloudflare’s models, the path our pilot called.
The method, in four lines. Write the $0 rule first and run it on every task. Send the same batches to one decision model. Log probability, evidence and reference for every decision. Buy decisions only where the rule loses. Teams who want the harness and the blind-review habit taught in-house get it through AI Enablement ($5,000).
Glossary
| Term | Meaning |
|---|---|
| Decision model | An AI model that takes a state plus typed questions and returns a probability for each allowed answer in one pass, with no generated text (TypeSafe docs; Cloudflare). |
| System One model | TypeSafe’s name for the category: “a class of AI models built to make fast, structured decisions that software can use directly”; also the request shape Clef and Liquid follow. |
| State | The evidence the model decides about: text or JSON for Jev; text, JSON, images and video for Clef. |
| Typed question (primitive) | A question with a declared answer type; the three shared primitives are noul, choice and score. |
| Noul | A yes/no primitive: is this statement true of the state? Returns a single value from 0 to 1 (TypeSafe docs). |
| Choice | Pick one option from a list you supply; returns the choice, a probability per option and a confidence. |
| Score | Place the state on an ordered rubric you define; returns the level, a legend and a probability per level. |
| Probability | The model’s output for each allowed answer; sums to one across the options of a question. |
| Confidence | One number from 0 to 1 that summarises how concentrated the probabilities are: 1.0 when all on one outcome, 0 when spread evenly (TypeSafe docs). |
| Calibration | Of all answers given at probability p, about a fraction p should be correct; measured across groups of predictions, not per answer (TypeSafe docs). |
| Threshold | The probability at or above which your code acts; TypeSafe recommends no automatic action below 0.5 and tiers that rise with the cost of a mistake. |
| Cascade (two-tier guardrail) | Run the cheap decision model on everything, escalate the band between thresholds to a frontier model, send the lowest band to a person. |
| Classifier | A model trained on labelled examples for a fixed label set; changing the labels means retraining. |
| JSON mode / structured outputs | An LLM constrained to emit valid JSON, optionally matching a schema; the output is still generated text (OpenAI docs). |
| DMN (the other sense) | Decision Model and Notation, the Object Management Group’s standard for specifying business decisions and rules by hand; unrelated to the AI sense. |
| Agreement vs accuracy (our term) | Agreement is the share of decisions matching post-hoc references; accuracy would need human-verified labels, which only grounding (RAGTruth), rules and leaks had. |
| Keyword baseline | A local rule with no API call, scored on the same batches; $0, 139/155 (89.7%) in our pilot, 12/12 on injection. |
Why Startrise publishes this
We re-measure whenever a new model tier appears, because routing decisions rot, and the model choice ships with its evidence. If you’re building the agents, that’s AI Agents That Act (from $3,500). If the agent is already live and needs the two-tier gate, that’s AI Monitoring & Guardrails (from $2,500). The study lives at Startrise AI Labs. When Jev or Clef ships a new version, we’ll run the same 168 decisions again and say what changed.
Questions we actually get
What is a decision model in AI?
A decision model is an AI model that takes a state (text, JSON, for some models images) plus a typed question with its allowed answers, and returns a probability for each answer in one pass. It generates no text. TypeSafe's Jev and Cloudflare's Clef are the two we have measured; both expose yes/no, multiple-choice and scored questions.
Is a decision model just an LLM that returns JSON?
No. An LLM asked for JSON still generates text token by token, and its "confidence" field is more generated text. In our pilot the two chat models returned 1.00 whether right or wrong and produced two output-cap truncations in 843 attempts; Jev and Clef returned graded probabilities and had zero errors across 336 calls.
Is a decision model the same as a classifier?
No, but it competes with one. A classic classifier is trained on labelled examples for a fixed label set; a decision model needs none and takes the label set from the request. Red Hat's October 2026 benchmark concluded that decision models like Jev "do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy."
What is a System One model?
TypeSafe's name for the category: "a class of AI models built to make fast, structured decisions that software can use directly," after Kahneman's fast System 1 thinking. Jev is the first. The phrase also names the request format: Cloudflare says Clef "follows the System One API," and Liquid AI's decision endpoint carries the same word.
Do decision models hallucinate?
They cannot invent prose, because they produce none, but they can be wrong, and sometimes confidently. On our injection task Jev and Clef scored 4/12 against a $0 keyword rule's 12/12 (recall-only). Clef Flash averaged 0.91 confidence when right and 0.85 when wrong (n=138/29), too narrow a gap for a threshold to catch.
How much does a decision model cost per decision?
Price decisions, not tokens: output is free on Jev, Clef and Clef Flash. Metered over 168 identical decisions each, Jev cost about $17 per million decisions, Clef Flash $24, Clef $63, GPT-6 Luna $27 and Claude Opus 5.5 $2,708. Rate cards: Jev $0.042, Clef Flash $0.09, Clef $0.24 per million input tokens.
Can you run a decision model yourself?
The Clef decision model and Clef Flash, yes: Cloudflare released the weights under Apache 2.0 on Hugging Face (open weights, not open source; the training pipeline is private). The Jev decision model is hosted and closed-weight. Red Hat has shown DiffusionGemma, another open decision model, running on vLLM. We measured Jev and Clef through a gateway, not self-hosted.
Is a decision model the same as DMN?
No. DMN, Decision Model and Notation, is the Object Management Group's standard for specifying business rules and decision tables by hand. The AI decision model described here is a trained model that returns probabilities over answers you declare. The phrase is shared; the objects are not, and this page covers the AI sense only.