+1 (415) 347-6981Get started
Fixed prices · You own the code · Proposal in 24h hello@startrise.io →

Jev vs Clef: 168 Decisions, 5 Models, One Pilot (2026)

Jev vs Clef on 168 identical decisions: Jev 149/167, Clef 150/167, Clef at 3.8x the price and 2x the latency. Where they split, and where a $0 baseline won.

Cloudflare shipped Clef and Clef Flash on October 1, 2026 as “fully Jev-API compatible” decision models. On October 4 Startrise ran both, plus Jev 1.13, GPT-6 Luna, Claude Opus 5.5 and a $0 keyword baseline, through ten small decision tasks with identical evidence and questions. Jev 1.13 agreed with the references on 149 of 167 reviewed decisions; Clef on 150 of 167. Clef cost 3.8x more per decision ($63 against $17 per million decisions) and, through the path we called, took twice as long (655 against 320 ms serialized per decision). One decision in 167 is noise, and agreement with post-hoc references is not accuracy.

Nobody else has this data. Every page ranking for “jev vs clef” today re-plots Cloudflare’s launch table; the one independent test we found ran on private data, no chat model or baseline in frame. The hands-on decision-model study is the only published dataset where all six rows got the same evidence and questions across ten tasks: 168 decisions per model, 843 attempts, $0.4769 spent, every record on a public evidence page. The study page is a recorded replay; the receipts under it are real timings and bills.

Key Takeaways

  • Jev 1.13 149/167 (89.2%) and Clef 150/167 (89.8%): one decision apart, inside noise, on agreement with post-hoc references
  • Clef cost 3.8x Jev per decision ($63 vs $17 per million) and ran at twice Jev’s serialized time through our gateway path (655 vs 320 ms); Cloudflare’s benches say the opposite
  • They split on two tasks: Clef took links 9/9 to 5/9; Jev took ads 15/18 to 13/18. On seven of ten they tied
  • A $0 keyword baseline scored 12/12 on injection (recall-only) where Jev and Clef scored 4/12 and Opus 5.5 8/12; it lost on feed, launches, ads and links
  • Claude Opus 5.5 scored 151/167 (90.4%) for $2,708 per million decisions: 162x Jev’s price for 2 more correct decisions

What are decision models, and why compare Jev and Clef?

A decision model takes a state (text or JSON, and for Clef also images) plus typed questions, and hands back a probability for each allowed answer in one pass. The Jev decision model has three primitives: a yes/no “noul”, a multiple choice and an ordered score. TypeSafe’s documentation says every question is evaluated in parallel inside a single call. No text comes out, so there’s nothing to parse and nothing to truncate. That is the decision model vs LLM difference in one line: a chat model prompted for JSON writes its answer as text, and text can run out of room. The full definition, the glossary and the disambiguation from DMN are in what a decision model is.

Jev 1.13 is TypeSafe AI’s hosted, closed-weight model. It launched on September 15, 2026 per InfoQ, costs $0.042 per million input tokens with output free, takes text only, and allows 64k tokens per request with 32k for the state. TypeSafe’s launch post calls it “193.6x faster, 444.6x cheaper” than frontier chat models. Vendor numbers on vendor evals; our meter answers them below.

The Clef decision model is Cloudflare’s answer. The launch post describes two models post-trained from frozen Qwen backbones, Clef (27B, Qwen3.8-27B) and Clef Flash (9B, Qwen3.5-9B), released under Apache 2.0. The weights are open, the training pipeline isn’t, so Hacker News settled on “open weights, not open source”. On Workers AI they run with a 65,536-token context, up to 64 questions per request and image input, at $0.24 and $0.09 per million input tokens, no output charge. Cloudflare’s benches put median latency at 209.3 ms for Clef, 38.8 ms for Clef Flash and 524.1 ms for Jev. Vendor numbers on vendor benches; ours sit beside them below.

The pitch is substitution. Cloudflare’s changelog says you switch a Jev integration to Clef by changing the endpoint and model. We logged the October 1 release in our model tracker’s Clef release entry, then priced the swap by sending both models the same requests. The chat-model routing guide covers the chat layer; this is the first Startrise AI Labs study on the layer underneath.

How did we test Jev vs Clef?

Ten tasks in four operations: rank (rentals, ads, links), filter (feed), label (bursts, launches) and highlight (injection, leaks, rules, grounding). Data came from Inside Airbnb Boston, the CHASM RedNote ad corpus, the Hacker News API, Bluesky posts from two accounts, Microsoft’s BIPIA attacks, RAGTruth, and Startrise’s own pages and fictional leak fixtures. Two or three presets per task, 26 in all, each a batch of 6 to 12 records.

Every model saw the same evidence and question per preset: at most 1,400 serialized request bytes including provider controls, no truncation, a 0.5 threshold. Six rows: Jev 1.13 (typesafe/jev-1.13), Clef (cloudflare/clef), Clef Flash (cloudflare/clef-flash), GPT-6 Luna and Claude Opus 5.5 as chat models prompted for JSON under a 250-token output cap, and a local keyword baseline with no API call. Model ids and rate cards are OpenRouter metadata, verified October 4. Plan jev-pilot-2026-10-04 finished at 23:06 UTC: 130 batches, 840 completed calls, 3 failures, $0.4769 actually spent.

References came after the run. Two blind AI reviewers who never see model outputs label each record. A reference counts only when both agree at high or medium confidence; disagreements leave the denominator and are counted in the open. Rules and leaks use deterministic checks, grounding uses RAGTruth’s human annotations, and injection uses BIPIA’s attack labels, where every record is an attack, so it measures recall only. Why we use two blind reviewers is our 839-call judge audit. The hands-on decision-model study replays every task; every record, decision, reference, cost and failure is on the evidence page.

Read this before the numbers

The study’s own warning goes first: “These small pilots do not establish general accuracy.”

Agreement is not accuracy. Except for grounding (human labels) and rules and leaks (deterministic), every score is agreement with two post-hoc AI reviewers.

Batches are tiny. Six to twelve records each, 167 reviewed decisions per API model, so the reviewed denominator sits beside every percentage. One decision is 0.6 points.

There is no single score. The ten tasks use unlike metrics and the study refuses to average them. The aggregate rows below sum correct over reviewed decisions: a reading aid, not THE score.

Some tasks are simulations. Ads persona matches are simulated judgments, not conversion evidence. The bursts snapshot covers two accounts, not the network. Injection is recall-only.

One timing trial. Latency and cost are recorded serialized batch time and billed cost including retries, through the gateway path we called. “Pilot only; no calibration tuning or held-out accuracy claim.”

Outside context: Red Hat Developers benchmarked Jev 1.13 against eight guardrail models on October 2, Clef not included, and concluded that “decision models like Jev do not reliably outperform LLM-as-a-judge, pre-trained predictive models, or open source decision models in speed or accuracy”.

Which is more accurate, Jev or Clef?

Neither, measurably. The aggregate first, labelled agreement, not accuracy:

ModelCorrect / reviewedAgreementFalse positivesFalse negatives
Jev 1.13149/16789.2%216
Clef150/16789.8%215
Clef Flash138/16782.6%227
GPT-6 Luna150/16789.8%314
Claude Opus 5.5151/16790.4%511
Keyword baseline139/15589.7%115

Sum of reviewed decisions across ten tasks with unlike metrics. The baseline has no grounding run, so its denominator is 155. Per-task results with cost and time are at the study’s results view; every record is on the evidence page.

The per-task table is where the comparison lives. Every cell is correct/reviewed.

Task (operation, reviewed)Jev 1.13ClefClef FlashGPT-6 LunaClaude Opus 5.5Keyword baseline
rentals (rank, 36)36/36 (100%)36/36 (100%)36/36 (100%)36/36 (100%)36/36 (100%)36/36 (100%)
feed (filter, 18)17/18 (94%)17/18 (94%)15/18 (83%)17/18 (94%)17/18 (94%)15/18 (83%)
bursts (label, 26)25/26 (96%)25/26 (96%)26/26 (100%)26/26 (100%)22/26 (85%)25/26 (96%)
launches (label, 18)18/18 (100%)17/18 (94%)18/18 (100%)17/18 (94%)17/18 (94%)15/18 (83%)
ads (rank, 18)15/18 (83%)13/18 (72%)13/18 (72%)13/18 (72%)14/18 (78%)12/18 (67%)
links (rank, 9)5/9 (56%)9/9 (100%)3/9 (33%)8/9 (89%)8/9 (89%)7/9 (78%)
injection (highlight, 12, attacks only)4/12 (33%)4/12 (33%)3/12 (25%)4/12 (33%)8/12 (67%)12/12 (100%)
leaks (highlight, 6)6/6 (100%)6/6 (100%)6/6 (100%)6/6 (100%)6/6 (100%)6/6 (100%)
rules (highlight, 12)11/12 (92%)11/12 (92%)9/12 (75%)12/12 (100%)12/12 (100%)11/12 (92%)
grounding (highlight, 12, human labels)12/12 (100%)12/12 (100%)9/12 (75%)11/12 (92%)11/12 (92%)not run

Source: Startrise decision-model pilot, plan jev-pilot-2026-10-04. Agreement with post-hoc references; rules and leaks deterministic; grounding against RAGTruth human annotations; injection is recall on known attacks.

Jev and Clef tie on seven of ten tasks: rentals 36/36, feed 17/18, bursts 25/26, injection 4/12, leaks 6/6, rules 11/12 and grounding 12/12. Launches went to Jev by one (18/18 against 17/18), ads to Jev by two (15/18 against 13/18), and links to Clef by four (9/9 against 5/9), the widest per-task gap between them.

Clef Flash is the row to watch, not buy. It disagreed with Clef on 16 of 168 decisions and Clef was right on 15; Flash’s misses cluster in links (3/9), grounding (9/12) and rules (9/12).

Injection, where no decision model was any good (4/12 for Jev and Clef, 3/12 for Clef Flash), gets its own section below.

Which is faster, Jev or Clef?

Through the path we called, Jev. Serialized time per decision (recorded batch time divided by 168 decisions): Clef Flash 316 ms, Jev 320 ms, Clef 655 ms, GPT-6 Luna 1,424 ms, Claude Opus 5.5 3,221 ms. Average time-to-first-byte agrees: Clef Flash 314 ms, Jev 318 ms, Luna 545 ms, Clef 652 ms, Opus 1,946 ms.

Serialized ms per decision (168 decisions, one timing trial)Clef FlashJev 1.13ClefGPT-6 LunaClaude Opus 5.5316 ms (TTFB 314 ms)320 ms (TTFB 318 ms)655 ms (TTFB 652 ms)1,424 ms (TTFB 545 ms)3,221 ms (TTFB 1,946 ms)Blue: decision models. Orange: chat models prompted for JSON.
Source: Startrise decision-model pilot, plan jev-pilot-2026-10-04. One timing trial; serialized batch time divided by 168 decisions; not a forward-pass benchmark and not the study page’s animation speed.

Cloudflare’s table says the opposite. Both can be true. Cloudflare’s medians (Clef 209.3 ms, Clef Flash 38.8 ms, Jev 524.1 ms) are forward passes, vendor-measured, on Cloudflare’s own infrastructure. Ours are end-to-end through the gateway path we called, one trial, a 27B model arriving slower than a hosted proprietary one. Different distances. TypeSafe’s vendor-measured “193.6x faster” shrinks the same way: on our path Jev’s end-to-end ratio was about 4.5x against Luna and 10x against Opus 5.5.

If you call these through a gateway rather than Workers AI directly, time your own path before trusting either table. If speed is the constraint, Clef Flash’s 316 ms matches Jev’s 320 ms. It gives up 11 decisions for that.

Which is cheaper, Jev or Clef?

Jev, by 3.8x per decision on the meter. Rate cards first, per million tokens, from OpenRouter metadata verified October 4: Jev $0.042 in, $0 out; Clef $0.24 in, $0 out; Clef Flash $0.09 in, $0 out; GPT-6 Luna $0.10 in, $0.50 out; Claude Opus 5.5 $4 in, $20 out. Then the meter for 168 decisions each: Jev $0.0028, about $17 per million decisions; Clef Flash $0.0040, about $24 per million; Luna $0.0045, about $27 per million; Clef $0.0106, about $63 per million; Opus 5.5 $0.4550, about $2,708 per million.

Agreement vs cost per million decisions (log cost axis)92%88%84%80%$10$100$1,000$10,000Cost per million decisions, as metered (log scale)$0 keyword baseline: 139/155 (89.7%)Jev 1.13: $17, 149/167 (89.2%)Clef Flash: $24, 138/167 (82.6%)GPT-6 Luna: $27, 150/167 (89.8%)Clef: $63, 150/167 (89.8%)Claude Opus 5.5: $2,708, 151/167 (90.4%)Blue: decision models. Orange: chat models prompted for JSON.
Source: Startrise decision-model pilot, plan jev-pilot-2026-10-04. The y-axis is agreement with post-hoc references, not accuracy. The $0 baseline cannot sit on a log axis, so it is drawn as a reference line. The baseline has no grounding run, so its denominator is 155. One decision is 0.6 points.

Clef is 3.8x Jev per decision. Opus 5.5 is 162x Jev and 43x Clef, bought 2 more correct decisions than Jev and 1 more than Clef, and swallowed $0.4550 of the pilot’s $0.4769.

Token shape is why the card and the meter disagree. On the card, Clef’s $0.24 is 5.7x Jev’s $0.042. On the meter it came out at 3.8x, because Clef counted fewer input tokens for the same evidence (44,124 against Jev’s 66,605). Clef reports 0 output tokens; Jev reports about 20 per call, priced at $0. Both are input-only in practice. Opus 5.5 isn’t: its 8,068 output tokens at $20 per million are about a third of its $0.4550 bill.

One cross-check lands in range: a commenter on the Clef thread ran the rate-card arithmetic at 300 tokens per call and got about $12.60 per million decisions on Jev and $72 on Clef. Ours differ because the two services counted the same evidence differently: Jev averaged 396 input tokens per decision (66,605 over 168), Clef 263 (44,124 over 168). And TypeSafe’s “444.6x cheaper”, measured against an average of GPT-6 Astra and Fable 5.1, becomes 162x against Opus 5.5 on our meter.

That’s also the answer to “cheapest model for classification”: price per decision, not per token, because output is free. We print the bill next to the score for the same reason the Startrise LLM Benchmark does.

Where did Jev and Clef disagree?

Twelve of 168 shared decisions. Of the eleven with a resolved reference, Jev was right on 5 and Clef on 6; the twelfth is unresolved because the reviewers disagreed. Two are worth telling.

Links, Clef’s win. The task ranks Startrise’s own pages as next-step links for an article. Clef 9/9, Jev 5/9. In both the AI-tools and developer-tools presets, Jev rated “AI Labs” at p=0.44 and “AI enablement” at p=0.43 to 0.48, just under the line, and left them out; Clef rated them p=0.69 to 0.95 and kept them. The references said include: four decisions, one direction, all within 0.07 of Jev’s threshold.

The Connect a page experiment on the Startrise Jev study page, AI tools preset, Jev selected in the model switcher: the article node labelled AI tools on the left connects to three Startrise page cards on the right; AI monitoring and guardrails is highlighted as Useful next step, while AI Labs and AI enablement read Not selected
The links experiment (Connect a page), AI tools preset, Jev selected: Jev connects the article to “AI monitoring and guardrails” and leaves “AI Labs” (p=0.44) and “AI enablement” (p=0.43) unselected. Switch the model to Clef on the study page and both cards turn on (p=0.95 and p=0.86). Links overall: Clef 9/9, Jev 5/9. Recorded replay, not inference time. Open the study.

Ads, Jev’s win. The task matches RedNote ad copy to a persona. Jev 15/18, Clef 13/18, and the whole gap is the Busy operator preset. Jev kept “Make room on the dressing table” at p=0.52 and “A Tagi bag at half price” at p=0.50, right on its threshold; Clef dropped them at p=0.20 and p=0.13. The reference said keep.

The Match an audience experiment on the Startrise Jev study page, Busy operator preset, Jev selected: a RedNote ad titled A guide to robot vacuums on the left; on the right the audience Busy operator, the verdict Jev says it fits, and a row headed Would the other models choose it? showing Jev Fits (highlighted), Clef Fits, Clef Flash Fits, GPT-6 Luna Skip and Opus 5.5 Fits
The ads experiment (Match an audience), Busy operator preset, Jev selected: on “A guide to robot vacuums”, Jev, Clef, Clef Flash and Opus 5.5 all say it fits and GPT-6 Luna skips it. The two ads that split Jev and Clef in this preset, “Make room on the dressing table” and “A Tagi bag at half price”, are in the same deck. Persona fit does not predict real customer preference; these are simulated judgments, not conversion evidence. Recorded replay, not inference time. Open the study.

They shared errors too. Both made the same false positive on the feed Developer-tools preset, marking “Remove and Disable Apple macOS 27 AI Models Tool” as developer-relevant when the reviewers said no. Opus 5.5 didn’t. Both missed the BIPIA “Task Automation-1” attack on the strict scan; on the balanced scan Clef caught it (p=0.83) while Jev caught Business Intelligence-0 (p=0.53) instead.

Those probabilities are the pilot’s operational finding. Jev’s average confidence in its chosen side was 0.86 when right and 0.69 when wrong (n=149 and 18), the widest gap of any model. Its mistakes live in its low-confidence calls, so a stricter threshold would catch some of them and cost you recall. Clef’s gap is narrower, 0.94 against 0.84 (n=150 and 17). Clef Flash is confidently wrong more often, 0.91 against 0.85 (n=138 and 29). GPT-6 Luna and Claude Opus 5.5 returned 1.00 whether right or wrong (n=150/17 and 151/16); a chat model prompted for JSON hands you a verdict and nothing to threshold on. We didn’t tune thresholds. The study’s phrase is “pilot only; no calibration tuning.”

Where did a $0 keyword baseline beat both?

Injection. The keyword baseline scored 12/12 on BIPIA’s embedded-instruction attacks because the attack strings carry literal instruction phrases a regex catches. Jev and Clef scored 4/12, Clef Flash 3/12, GPT-6 Luna 4/12, Claude Opus 5.5 8/12. Every record is an attack, so this is recall-only. The baseline’s false-positive rate went unmeasured; a filter that flags everything also scores 12/12.

Two other tasks separated nobody: rentals was 36/36 and leaks 6/6 for every row, baseline included.

The baseline’s aggregate is 139/155 (89.7%) against Jev’s 149/167 (89.2%) and Clef’s 150/167 (89.8%), on a smaller denominator because it never ran grounding. Don’t read that as “use regex.” The paid models earn their keep where the baseline fails. Feed: baseline 15/18, Jev and Clef 17/18. Launches: 15/18 against Jev’s 18/18. Ads: 12/18 against Jev’s 15/18. Links: 7/9 against Clef’s 9/9. Those four tasks are the purchase. The other six came free.

So run the $0 baseline on every task first, then buy decisions only where it loses. That ordering is the first thing we build in a workflow automation engagement.

When should you still pay for Claude Opus 5.5?

On this data, for one task class. Injection: Opus 5.5 scored 8/12 against 4/12 for Jev and Clef. Everywhere else you paid 162x Jev’s price for 2 more correct decisions over 167, and it gave ground on bursts (22/26 against 25/26, with 4 false positives) and ads (14/18 against Jev’s 15/18).

The failure log is the other half. Three failures in 843 attempts, all on chat models. On feed / AI tools, Opus 5.5 returned HTTP 402, a provider credit pre-check, retried and settled. On feed / Developer tools, GPT-6 Luna’s reasoning ate 150 of a 250-token output cap and left no visible decision. On ads / Busy operator, Opus 5.5 truncated its structured JSON at the 250-token output ceiling. Jev and Clef: zero errors, zero retries across 336 calls, because there’s no text to truncate.

A chat model asked for a decision writes a document with the decision inside, and documents run out of room; the 128k output-ceiling finding from the day before is the same failure at scale. Decision models can be wrong, as the injection row shows. But they always answer.

The Benchmark results view on the Startrise Jev study page for the Check a fact experiment, Unsupported claims preset: Jev 6 of 6 human-label agreement, $0.000150 full batch cost, 1.80 seconds; Clef 6 of 6, $0.000656, 4.43 seconds; Clef Flash 4 of 6, $0.000246, 4.05 seconds; more model rows continue below the crop
The results view, Check a fact (grounding) on the Unsupported claims preset: Jev 6/6 for $0.000150 in 1.80 seconds, Clef 6/6 for $0.000656 in 4.43 seconds, Clef Flash 4/6 for $0.000246 in 4.05 seconds. The reviewed denominator sits on every row, with billed cost and serialized time beside it. Recorded replay, not inference time. Open the study.

The guardrail design writes itself: run the cheap decision model first, escalate to the frontier model only on low confidence. Jev’s 0.86-against-0.69 gap makes that possible; a chat model’s flat 1.00 doesn’t. When we add AI Monitoring & Guardrails to an agent in production, from $2,500, that two-tier gate is the shape it takes.

Jev or Clef: how to pick

A routing verdict on this data only; no crown.

  • Default for text-only decisions at volume: Jev 1.13. Tied on agreement (149/167 against 150/167), 3.8x cheaper at $17 per million decisions, half Clef’s serialized time at 320 ms, and the only model whose confidence separated right from wrong by 0.17.
  • For images in the state, a 64k context, open weights, or a drop-in for existing Jev code: Cloudflare Clef. Budget 3.8x the price at $63 per million, time it on your own path, and expect a different answer on about one decision in fourteen (12 of 168 here).
  • When speed matters more than the last seven points of agreement: Clef Flash. 316 ms against Jev’s 320, finished 138/167, with misses in links (3/9), grounding (9/12) and rules (9/12).
  • Before any of them: the keyword baseline, on every task. Where it loses is the list of decisions worth paying for.
  • Only where baseline and decision model both fail (here, injection): Claude Opus 5.5. 8/12 against 4/12, at $2,708 per million decisions and 3,221 ms, with a retry for the 250-token ceiling wired first.
  • Not yet: OpenAI’s Decision API built on Luna, reported by The New Stack on September 29, and Cloudflare’s RL fine-tuning service for Clef. We tested neither.

Whichever model tops that list, the low-confidence branch usually ends at a person; approval workflows for agents and human-in-the-loop over MCP pick up where the probability leaves off.

Why we ran this pilot

Startrise re-measures every time a new model tier shows up, because routing decisions rot. Cloudflare shipped Clef on October 1. By October 4 every decision in this pilot had a recorded probability, a reference label, a cost and a timing on a public evidence page, failures included. That discipline is what clients buy when we build AI Agents That Act: agents on Mastra with routing, fallbacks, guardrails and approval gates, from $3,500, where the model choice ships with its evidence.

AI Content Pipelines ($5,000) covers the feed, launches and bursts class of task with a decision layer in front of human QA gates. Teams who’d rather run their own pilots get the harness and the blind-review habit from AI Enablement ($5,000). For frontends rather than decisions, our frontend buying guide is the companion piece.

The study is live at the hands-on decision-model study. When Jev or Clef ships a new version, we’ll run the same 168 decisions again and say what changed.

Questions we actually get

Is Clef more accurate than Jev?

Not measurably in the Startrise pilot. Across 167 reviewed decisions each, Jev 1.13 agreed with the references on 149 and Clef on 150. They disagreed on 12 of 168 shared decisions; of the 11 resolved, Jev was right 5 times and Clef 6. The split is by task: Clef won links 9/9 to 5/9, Jev won ads 15/18 to 13/18.

Which is cheaper per decision, Jev or Clef?

Jev. On 168 identical decisions Jev billed $0.0028, about $17 per million decisions, and Clef $0.0106, about $63 per million, so Clef cost 3.8x. Both are input-only in practice: Clef reports zero output tokens and Jev's roughly 20 output tokens per call are priced at $0. Claude Opus 5.5 cost $2,708 per million decisions on the same work.

Is Clef faster than Jev?

Not through the path we called. Clef averaged 655 ms serialized per decision and 652 ms to first byte; Jev 320 ms and 318 ms. Cloudflare's published medians (Clef 209.3 ms, Jev 524.1 ms) measure the forward pass on its own infrastructure. Clef Flash matched Jev's speed at 316 ms but agreed on 11 fewer decisions, 138/167 against 149/167.

What is a decision model, and how is it different from an LLM?

A decision model takes a state plus typed questions and returns a probability for each allowed answer in one pass; it generates no text. A chat model prompted for JSON generates tokens that can truncate or drift. Jev and Clef had zero errors across 336 pilot calls; the two chat models produced three failures, including two 250-token output-ceiling truncations.

Can a keyword filter replace Jev or Clef?

Sometimes; check first. Our $0 keyword baseline scored 12/12 on injection, where Jev and Clef scored 4/12 and Opus 5.5 8/12, because the test attacks contain literal instruction phrases. It matched every model on rentals and leaks. It lost on feed (15/18), launches (15/18), ads (12/18) and links (7/9), which is where paying for decisions earned its keep.

Is Clef a drop-in replacement for Jev?

Cloudflare says Clef is fully Jev-API compatible and that you switch by changing the endpoint and model. Our pilot sent identical requests to both without code changes. Expect different behaviour, though: on 168 shared decisions they disagreed 12 times, and Clef rated two Startrise pages at p=0.69 to 0.95 where Jev sat at p=0.43 to 0.48.

Next step

Want this applied to your product?

hello@startrise.io
Most projects start with a 24-hour proposal. Get a proposal in 24h →