Cloudflare brought a model to the Jev fight
Cloudflare launched its own decision models, Clef and Clef-flash, to take on TypeSafe’s Jev, and it put real weight behind them: a launch post, open weights, a Workers AI endpoint, and a claim that Clef beats Jev “in 3 out of 4 areas.” I took the claim to my own labelled data. Is it good enough to switch?
Short version: Clef ties Jev on accuracy, and the paired test says the gap is noise. It abstained more on nonsense input than Jev in a small probe set, and it calibrates better on one task. It also costs 4.2x what Jev costs per decision and answers 2.6x slower at p50. Clef-flash is cheaper than Clef, loses to Jev on accuracy, and still costs 1.6x Jev. The “3 of 4” lead did not reproduce on private data, and the “fully API-compatible” claim has a real hole. Jev stays my default.
Full example: Clone the working files at github.com/KingPin/sumguy-examples/llm/clef-vs-jev
This is round three. The first two rounds scored the self-hosted clones: Jev vs Von vs SemIf: Real Numbers and The Update That Broke Confidence. This round is hosted only: Jev against Clef against Clef-flash. No self-hosted systems this time.
What Cloudflare is selling
Per Cloudflare’s launch post (fetched 2026-10-01), Clef has a vision encoder, so the state can include images. Jev is text-only. Clef has a 64k context window (the model page says 65,536 tokens) “compared to Jev’s 32k”. Cloudflare calls it “fully API-compatible, so you can make the swap extremely easily.”
Under the hood, Clef is a frozen Qwen3.8-27B with a routing head and a rank-256 LoRA. Clef-flash uses the same recipe on a frozen Qwen3.5-9B. The weights are open under Apache 2.0 on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash. Cloudflare says it doesn’t read, store, or train on requests, and it offers fine-tuning through a forward-deployed-engineer team, with self-serve promised later.
The “3 of 4” claim comes from TypeSafe’s own WorkflowEvals suite, which Cloudflare ran itself (invoice processing, for example: Clef 64.7, Clef-flash 57.1). They also cite the Jev Decision Index. Those are public eval suites. Public suites are also where the vendor controls the setup.
The three primitives are the same as before. choice returns a choice, probabilities, and a confidence. noul returns a single probability. score takes ordered levels and returns a fractional score plus a confidence.
One gotcha before any benchmarking: Clef’s confidence is not the top probability. In one smoke test the top probability was 0.971 and the confidence came back as 0.887. Threshold on the field you mean.
The setup
Same method as the earlier rounds, short version. The labels came from the git history of Wyrmhole, the private Discord RPG bot named in the earlier posts. The repo is private, so there are no commit messages or file paths here. I rebuilt the corpus with a fixed seed at a 2,750-commit snapshot and stripped conventional-commit prefixes from the inputs, because otherwise you grade the model on its own answer key.
The same 440 items went to all three models with the same question sets. Every task has a free dumb baseline scored alongside. The t3 and t4 definitions differ from earlier rounds and this is a fresh draw, so do not compare numbers across posts.
- t1 commit type:
choice, 6 options (feat, fix, docs, chore, test, refactor), n=90, 15 per type. State is the subject plus diffstat plus up to 3.5k of patch. - t2 file routing:
choice, 12 candidate files (11 random decoys plus the true file, each described by its module docstring), n=80. Random guessing gets 8.3%. Option keys are dotted paths (app.auth.py) for all three models, because of the slash bug covered below. - t3 is-a-feature:
noul, n=180 (30 features, 150 not). State is the subject plus the changed-file list. - t4 change breadth:
score, 3 levels (one file, two to four, five or more), n=90, 30 per level. State is the subject only.
The endpoints, trimmed to what matters. Jev goes through OpenRouter at https://openrouter.ai/api/alpha/decisions with the model ~typesafe/jev-latest, which resolved to typesafe/jev-1.13-20260917 during testing. Clef goes through Workers AI at https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef, and Clef-flash at the same path ending clef-flash. The body is {"model":"clef"|"clef-flash","state":...,"questions":...}.
The question file for t1 looks like this (it is questions.json in the companion repo):
{"type":{"type":"choice","instructions":"Which conventional-commit type best describes this commit?","criteria":{"feat":"adds a new feature or capability","fix":"fixes a bug","docs":"documentation only","chore":"maintenance, tooling, config, or dependencies","test":"adds or changes tests only","refactor":"restructures code without changing behavior"}}}And ask.sh sends one state to whichever providers have credentials set:
export JEV_API_KEY=sk-or-...export CLOUDFLARE_ACCOUNT_ID=...export CLOUDFLARE_API_TOKEN=...
./ask.sh "handle empty response from the weather API"On 2026-10-01 that state, with questions.json, gave this. Clef answered fix with confidence 0.4544 on 227 input tokens. Clef-flash answered fix with confidence 0.8399, also 227 input tokens. Same answer, very different certainty. Keep that in mind for the gate results below.
The results
Same 440 rows, 2026-10-01. Latency was measured from a single home machine, so treat the gaps as the signal and the raw milliseconds as local.
| task | metric | Jev | Clef | Clef-flash | free baseline |
|---|---|---|---|---|---|
| t1 commit type | accuracy | 80.0% | 82.2% | 73.3% | 48.9% (path rule) |
| t1 | accuracy at ≥0.85 confidence (coverage) | 90.5% (70%) | 100% (31%) | 100% (13%) | |
| t2 file routing | accuracy | 80.0% | 81.2% | 76.2% | 47.5% (keyword overlap) |
| t2 | accuracy at ≥0.85 (coverage) | 97.7% (55%) | 100% (30%) | 95.8% (30%) | |
| t3 is-a-feature | AUC | 0.889 | 0.870 | 0.784 | 0.565 (feature-word keyword) |
| t3 | Brier (lower is better) | 0.119 | 0.098 | 0.121 | 0.139 (base rate) |
| t3 | accuracy at 0.5 | 82.8% | 87.8% | 80.6% | 83.3% (always say no) |
| t4 breadth | Spearman | 0.240 | 0.003 | 0.050 | 0.225 (word count) |
| t4 | exact / MAE | 43.3% / 0.677 | 34.4% / 0.771 | 37.8% / 0.700 | 33.3% (constant) |
| latency | p50 / p95 ms | 231 / 293 | 609 / 1,102 | 448 / 827 | |
| tokens | input, 440 calls | 271,189 | 200,317 | 200,317 | |
| cost | per 1,000 decisions | $0.026 | $0.109 | $0.041 |
Why “ahead by 2 points” is a tie
On t1 Clef scores 82.2% against Jev’s 80.0%. That looks like a win until you pair the rows. I ran an exact McNemar test against Jev on the same items. On t1, Clef was right where Jev was wrong 5 times, and Jev was right where Clef was wrong 3 times (p=0.73). Clef-flash split 4 against 10 (p=0.18). On t2, Clef split 5 against 4 (p=1.00) and Clef-flash split 3 against 6 (p=0.51).
So Clef’s lead is 2 items on t1 and 1 item on t2. That is a tie. Clef and Clef-flash give the same answer as Jev on 81 to 88% of rows, so most of the time you pay four times as much for the same verdict.
The confidence gate
Clef’s gate is conservative. Only 31% of t1 rows clear 0.85, against 70% for Jev, but every one of those 28 rows is right. Jev answers more than twice as many rows confidently and still holds 90.5% accuracy on them.
Which one you want depends on what a wrong auto-decision costs against what a human review costs. My opinion: if the output triggers an action with no review, Clef’s gate is better, because 28 of 28 beats 57 of 63 when nobody is checking (small samples, so treat it as a lean). If you care about throughput and a human catches the stragglers, Jev’s gate is better.
t3 is a split decision
Jev ranks better (AUC 0.889 against 0.870). Clef calibrates better (Brier 0.098 against 0.119). Clef’s 87.8% accuracy at 0.5 beats the always-say-no baseline of 83.3%, while Jev’s 82.8% does not. Clef-flash loses on all three t3 metrics.
t4 is weak for everyone
Subject-only state does not carry how many files changed. Clef and Clef-flash sit at or below the word-count heuristic (Spearman 0.225), and Jev barely clears it at 0.240. When the free baseline ties everyone, the task needs more state, and a pricier model will not supply it.
Junk in, what comes out
I built four junk states: a single ., lorem ipsum, “what is the weather in Lisbon”, and the digits 4829173650. Each went against the 6-option commit-type question and a 12-option routing question, so 8 probes per model. I counted how many cleared the 0.85 gate.
Clef: 0 of 8. Clef-flash: 0 of 8. Jev: 2 of 8. The digits got chore at 0.89. The weather question got routed to the bot’s chat module at 0.90, which is arguably defensible, since that module is the chat bot. No model collapsed every junk input to the same answer, which is more than the previous round could say about Laya. This is the part where Clef earns its price, if abstention is what you buy it for.
The price, as of October 2026
Workers AI bills in neurons at $0.011 per 1,000 neurons. Clef costs $0.240 per million input tokens (21,818 neurons per million). Clef-flash costs $0.090 per million (8,182 neurons per million). The pricing page, updated October 1, 2026, lists no output price. Jev costs $0.042 per million input tokens with output free. The TypeSafe docs and the OpenRouter listing agree, and Jev Retagged 840 Posts For 26 Cents covers the details, including the $5 credit that TypeSafe’s billing panel labels monthly.
The sticker ratio is 5.7x (0.24 over 0.042). The real gap is smaller, because Clef’s tokenizer counted 26% fewer tokens for identical requests: 455 tokens per call against Jev’s 616. That brings the per-decision gap to 4.2x for Clef and 1.6x for Clef-flash.
Measured cost per 1,000 decisions: Jev $0.026, Clef $0.109, Clef-flash $0.041. Per million decisions that is about $26, $109, and $41. At bulk scale, Clef costs you an extra $83 per million decisions for a tie.
The full benchmark round, 1,320 calls across three models plus the probes, used about 6,100 Cloudflare neurons. That fits inside one day’s free tier. Total spend: Jev $0.0114, Clef $0.0481, Clef-flash $0.0180. The Clef numbers are what it would bill. The actual bill was $0.
The free tier changes the hobby math
Workers Free and Workers Paid both include 10,000 neurons per day, resetting at 00:00 UTC. That is about 458,000 Clef input tokens, or roughly 1,000 decisions per day at this benchmark’s 455 tokens per call. Clef-flash stretches to about 1.22M tokens, or roughly 2,700 decisions a day. Beyond the free allocation you need Workers Paid, which has a $5 per month minimum charge per account (Workers pricing page, updated August 28, 2026). On the Free plan, going over the limit makes requests fail with an error.
At hobby scale, price is irrelevant. A thousand free Clef decisions a day covers a lot of side projects.
The “fully API-compatible” gotcha
Cloudflare says you can swap Clef in for Jev easily. Two differences say otherwise.
First, the response wrapper. Workers AI returns {"result":{model,answers,usage},"success":...,"errors":[]}. Jev returns {model, answers, usage} unwrapped. Your client has to unwrap result. That is small, and it already breaks a literal drop-in swap.
Second, and worse: Workers AI rejects a choice question with HTTP 400 when every option key contains a /. One slash-free key makes the same request pass. Both Clef models do this. Jev accepts the identical request. The error message misleads you, too:
AiError: Bad input: Error: required properties at '/' are 'model,state,questions'That arrives with code 5006, even though all three properties are present. File paths as option keys are the obvious real-world use for file routing, so this hits the exact “swap it in” scenario Cloudflare advertises. The repro from slash-key-repro.sh sends three requests to Clef-flash:
export CLOUDFLARE_ACCOUNT_ID=...export CLOUDFLARE_API_TOKEN=...
./slash-key-repro.shOutput, trimmed, verified again on 2026-10-01 with the companion script:
== all keys contain a slashHTTP 400[{"message":"AiError: Bad input: Error: required properties at '/' are 'model,state,questions' (...)","code":5006}]== one key without a slashHTTP 200== workaround: dotted keysHTTP 200The workaround is to use dotted keys such as app.auth.py and keep the real path in the description text ("app.auth.py":"app/auth.py: login and sessions"). It works, but your 2 AM self will not enjoy finding this out from a 400 that says your properties are missing when they are sitting right there.
Who should use what
Default to Jev. Same accuracy, about a quarter of the price per decision, and 2.6x faster at p50. If your volume is real, that settles it.
Pick Clef when you need images in the state, more than 32k of state, or want everything inside an existing Cloudflare account and Workers binding. Pick it too if never auto-acting on junk matters more to you than price. The open weights also mean you can self-host later if the hosted bill ever bites.
Clef-flash is hard to recommend. It costs more than Jev and is less accurate. Its only edge is staying inside Cloudflare at a lower price than Clef.
Small volumes change nothing above, except that the free tier makes cost a non-issue. Use whichever model has the feature you need.
Vendor benchmarks deserve suspicion. Cloudflare’s numbers come from public eval suites, and on this one private repo the lead vanished. The cure is cheap: mine your own labels, add a free baseline, and run a paired test. The method is in Jev vs Von vs SemIf: Real Numbers. Do it before you wire a threshold into anything that acts on its own.
Common Questions
Is Cloudflare Clef cheaper than Jev?
No. Clef costs $0.240 per million input tokens against Jev’s $0.042. Clef’s tokenizer counts 26% fewer tokens, so the measured gap is 4.2x: $0.109 against $0.026 per 1,000 decisions. Clef-flash is cheaper than Clef but still costs about 1.6x Jev per decision.
Can I use Clef for free on Workers AI?
Yes, up to a limit. Both Workers Free and Workers Paid include 10,000 neurons per day, which covers roughly 1,000 Clef decisions at 455 tokens each. On the Free plan, requests fail once you pass the limit. Workers Paid bills overage and has a $5 monthly minimum.
Is Clef a drop-in replacement for Jev?
Not quite. Clef wraps its response in a result object, which your client must unwrap. Workers AI also returns HTTP 400 when every option key in a choice question contains a slash. Dotted keys fix the slash problem, with the real path kept in the description.
Can I self-host Cloudflare Clef?
Yes. The weights are Apache 2.0 on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash. Clef sits on a 27B base model, so it needs serious GPU memory. Clef-flash uses a 9B base, which asks for much less hardware. Neither was tested self-hosted in this benchmark.