Skip to content
Go back

Clef vs Jev: Cloudflare's 4x Tie

By KingPin 13 min read
Clef vs Jev: Cloudflare's 4x Tie
Contents

Cloudflare brought a model to the Jev fight

Cloudflare launched its own decision models, Clef and Clef-flash, to take on TypeSafe’s Jev, and it put real weight behind them: a launch post, open weights, a Workers AI endpoint, and a claim that Clef beats Jev “in 3 out of 4 areas.” I took the claim to my own labelled data. Is it good enough to switch?

Short version: Clef ties Jev on accuracy, and the paired test says the gap is noise. It abstained more on nonsense input than Jev in a small probe set, and it calibrates better on one task. It also costs 4.2x what Jev costs per decision and answers 2.6x slower at p50. Clef-flash is cheaper than Clef, loses to Jev on accuracy, and still costs 1.6x Jev. The “3 of 4” lead did not reproduce on private data, and the “fully API-compatible” claim has a real hole. Jev stays my default.

Full example: Clone the working files at github.com/KingPin/sumguy-examples/llm/clef-vs-jev

This is round three. The first two rounds scored the self-hosted clones: Jev vs Von vs SemIf: Real Numbers and The Update That Broke Confidence. This round is hosted only: Jev against Clef against Clef-flash. No self-hosted systems this time.

What Cloudflare is selling

Per Cloudflare’s launch post (fetched 2026-10-01), Clef has a vision encoder, so the state can include images. Jev is text-only. Clef has a 64k context window (the model page says 65,536 tokens) “compared to Jev’s 32k”. Cloudflare calls it “fully API-compatible, so you can make the swap extremely easily.”

Under the hood, Clef is a frozen Qwen3.8-27B with a routing head and a rank-256 LoRA. Clef-flash uses the same recipe on a frozen Qwen3.5-9B. The weights are open under Apache 2.0 on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash. Cloudflare says it doesn’t read, store, or train on requests, and it offers fine-tuning through a forward-deployed-engineer team, with self-serve promised later.

The “3 of 4” claim comes from TypeSafe’s own WorkflowEvals suite, which Cloudflare ran itself (invoice processing, for example: Clef 64.7, Clef-flash 57.1). They also cite the Jev Decision Index. Those are public eval suites. Public suites are also where the vendor controls the setup.

The three primitives are the same as before. choice returns a choice, probabilities, and a confidence. noul returns a single probability. score takes ordered levels and returns a fractional score plus a confidence.

One gotcha before any benchmarking: Clef’s confidence is not the top probability. In one smoke test the top probability was 0.971 and the confidence came back as 0.887. Threshold on the field you mean.

The setup

Same method as the earlier rounds, short version. The labels came from the git history of Wyrmhole, the private Discord RPG bot named in the earlier posts. The repo is private, so there are no commit messages or file paths here. I rebuilt the corpus with a fixed seed at a 2,750-commit snapshot and stripped conventional-commit prefixes from the inputs, because otherwise you grade the model on its own answer key.

The same 440 items went to all three models with the same question sets. Every task has a free dumb baseline scored alongside. The t3 and t4 definitions differ from earlier rounds and this is a fresh draw, so do not compare numbers across posts.

The endpoints, trimmed to what matters. Jev goes through OpenRouter at https://openrouter.ai/api/alpha/decisions with the model ~typesafe/jev-latest, which resolved to typesafe/jev-1.13-20260917 during testing. Clef goes through Workers AI at https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef, and Clef-flash at the same path ending clef-flash. The body is {"model":"clef"|"clef-flash","state":...,"questions":...}.

The question file for t1 looks like this (it is questions.json in the companion repo):

{"type":{"type":"choice","instructions":"Which conventional-commit type best describes this commit?","criteria":{"feat":"adds a new feature or capability","fix":"fixes a bug","docs":"documentation only","chore":"maintenance, tooling, config, or dependencies","test":"adds or changes tests only","refactor":"restructures code without changing behavior"}}}

And ask.sh sends one state to whichever providers have credentials set:

Terminal window
export JEV_API_KEY=sk-or-...
export CLOUDFLARE_ACCOUNT_ID=...
export CLOUDFLARE_API_TOKEN=...
./ask.sh "handle empty response from the weather API"

On 2026-10-01 that state, with questions.json, gave this. Clef answered fix with confidence 0.4544 on 227 input tokens. Clef-flash answered fix with confidence 0.8399, also 227 input tokens. Same answer, very different certainty. Keep that in mind for the gate results below.

The results

Same 440 rows, 2026-10-01. Latency was measured from a single home machine, so treat the gaps as the signal and the raw milliseconds as local.

taskmetricJevClefClef-flashfree baseline
t1 commit typeaccuracy80.0%82.2%73.3%48.9% (path rule)
t1accuracy at ≥0.85 confidence (coverage)90.5% (70%)100% (31%)100% (13%)
t2 file routingaccuracy80.0%81.2%76.2%47.5% (keyword overlap)
t2accuracy at ≥0.85 (coverage)97.7% (55%)100% (30%)95.8% (30%)
t3 is-a-featureAUC0.8890.8700.7840.565 (feature-word keyword)
t3Brier (lower is better)0.1190.0980.1210.139 (base rate)
t3accuracy at 0.582.8%87.8%80.6%83.3% (always say no)
t4 breadthSpearman0.2400.0030.0500.225 (word count)
t4exact / MAE43.3% / 0.67734.4% / 0.77137.8% / 0.70033.3% (constant)
latencyp50 / p95 ms231 / 293609 / 1,102448 / 827
tokensinput, 440 calls271,189200,317200,317
costper 1,000 decisions$0.026$0.109$0.041

Why “ahead by 2 points” is a tie

On t1 Clef scores 82.2% against Jev’s 80.0%. That looks like a win until you pair the rows. I ran an exact McNemar test against Jev on the same items. On t1, Clef was right where Jev was wrong 5 times, and Jev was right where Clef was wrong 3 times (p=0.73). Clef-flash split 4 against 10 (p=0.18). On t2, Clef split 5 against 4 (p=1.00) and Clef-flash split 3 against 6 (p=0.51).

So Clef’s lead is 2 items on t1 and 1 item on t2. That is a tie. Clef and Clef-flash give the same answer as Jev on 81 to 88% of rows, so most of the time you pay four times as much for the same verdict.

The confidence gate

Clef’s gate is conservative. Only 31% of t1 rows clear 0.85, against 70% for Jev, but every one of those 28 rows is right. Jev answers more than twice as many rows confidently and still holds 90.5% accuracy on them.

Which one you want depends on what a wrong auto-decision costs against what a human review costs. My opinion: if the output triggers an action with no review, Clef’s gate is better, because 28 of 28 beats 57 of 63 when nobody is checking (small samples, so treat it as a lean). If you care about throughput and a human catches the stragglers, Jev’s gate is better.

t3 is a split decision

Jev ranks better (AUC 0.889 against 0.870). Clef calibrates better (Brier 0.098 against 0.119). Clef’s 87.8% accuracy at 0.5 beats the always-say-no baseline of 83.3%, while Jev’s 82.8% does not. Clef-flash loses on all three t3 metrics.

t4 is weak for everyone

Subject-only state does not carry how many files changed. Clef and Clef-flash sit at or below the word-count heuristic (Spearman 0.225), and Jev barely clears it at 0.240. When the free baseline ties everyone, the task needs more state, and a pricier model will not supply it.

Junk in, what comes out

I built four junk states: a single ., lorem ipsum, “what is the weather in Lisbon”, and the digits 4829173650. Each went against the 6-option commit-type question and a 12-option routing question, so 8 probes per model. I counted how many cleared the 0.85 gate.

Clef: 0 of 8. Clef-flash: 0 of 8. Jev: 2 of 8. The digits got chore at 0.89. The weather question got routed to the bot’s chat module at 0.90, which is arguably defensible, since that module is the chat bot. No model collapsed every junk input to the same answer, which is more than the previous round could say about Laya. This is the part where Clef earns its price, if abstention is what you buy it for.

The price, as of October 2026

Workers AI bills in neurons at $0.011 per 1,000 neurons. Clef costs $0.240 per million input tokens (21,818 neurons per million). Clef-flash costs $0.090 per million (8,182 neurons per million). The pricing page, updated October 1, 2026, lists no output price. Jev costs $0.042 per million input tokens with output free. The TypeSafe docs and the OpenRouter listing agree, and Jev Retagged 840 Posts For 26 Cents covers the details, including the $5 credit that TypeSafe’s billing panel labels monthly.

The sticker ratio is 5.7x (0.24 over 0.042). The real gap is smaller, because Clef’s tokenizer counted 26% fewer tokens for identical requests: 455 tokens per call against Jev’s 616. That brings the per-decision gap to 4.2x for Clef and 1.6x for Clef-flash.

Measured cost per 1,000 decisions: Jev $0.026, Clef $0.109, Clef-flash $0.041. Per million decisions that is about $26, $109, and $41. At bulk scale, Clef costs you an extra $83 per million decisions for a tie.

The full benchmark round, 1,320 calls across three models plus the probes, used about 6,100 Cloudflare neurons. That fits inside one day’s free tier. Total spend: Jev $0.0114, Clef $0.0481, Clef-flash $0.0180. The Clef numbers are what it would bill. The actual bill was $0.

The free tier changes the hobby math

Workers Free and Workers Paid both include 10,000 neurons per day, resetting at 00:00 UTC. That is about 458,000 Clef input tokens, or roughly 1,000 decisions per day at this benchmark’s 455 tokens per call. Clef-flash stretches to about 1.22M tokens, or roughly 2,700 decisions a day. Beyond the free allocation you need Workers Paid, which has a $5 per month minimum charge per account (Workers pricing page, updated August 28, 2026). On the Free plan, going over the limit makes requests fail with an error.

At hobby scale, price is irrelevant. A thousand free Clef decisions a day covers a lot of side projects.

The “fully API-compatible” gotcha

Cloudflare says you can swap Clef in for Jev easily. Two differences say otherwise.

First, the response wrapper. Workers AI returns {"result":{model,answers,usage},"success":...,"errors":[]}. Jev returns {model, answers, usage} unwrapped. Your client has to unwrap result. That is small, and it already breaks a literal drop-in swap.

Second, and worse: Workers AI rejects a choice question with HTTP 400 when every option key contains a /. One slash-free key makes the same request pass. Both Clef models do this. Jev accepts the identical request. The error message misleads you, too:

AiError: Bad input: Error: required properties at '/' are 'model,state,questions'

That arrives with code 5006, even though all three properties are present. File paths as option keys are the obvious real-world use for file routing, so this hits the exact “swap it in” scenario Cloudflare advertises. The repro from slash-key-repro.sh sends three requests to Clef-flash:

Terminal window
export CLOUDFLARE_ACCOUNT_ID=...
export CLOUDFLARE_API_TOKEN=...
./slash-key-repro.sh

Output, trimmed, verified again on 2026-10-01 with the companion script:

== all keys contain a slash
HTTP 400
[{"message":"AiError: Bad input: Error: required properties at '/' are 'model,state,questions' (...)","code":5006}]
== one key without a slash
HTTP 200
== workaround: dotted keys
HTTP 200

The workaround is to use dotted keys such as app.auth.py and keep the real path in the description text ("app.auth.py":"app/auth.py: login and sessions"). It works, but your 2 AM self will not enjoy finding this out from a 400 that says your properties are missing when they are sitting right there.

Who should use what

Default to Jev. Same accuracy, about a quarter of the price per decision, and 2.6x faster at p50. If your volume is real, that settles it.

Pick Clef when you need images in the state, more than 32k of state, or want everything inside an existing Cloudflare account and Workers binding. Pick it too if never auto-acting on junk matters more to you than price. The open weights also mean you can self-host later if the hosted bill ever bites.

Clef-flash is hard to recommend. It costs more than Jev and is less accurate. Its only edge is staying inside Cloudflare at a lower price than Clef.

Small volumes change nothing above, except that the free tier makes cost a non-issue. Use whichever model has the feature you need.

Vendor benchmarks deserve suspicion. Cloudflare’s numbers come from public eval suites, and on this one private repo the lead vanished. The cure is cheap: mine your own labels, add a free baseline, and run a paired test. The method is in Jev vs Von vs SemIf: Real Numbers. Do it before you wire a threshold into anything that acts on its own.

Common Questions

Is Cloudflare Clef cheaper than Jev?

No. Clef costs $0.240 per million input tokens against Jev’s $0.042. Clef’s tokenizer counts 26% fewer tokens, so the measured gap is 4.2x: $0.109 against $0.026 per 1,000 decisions. Clef-flash is cheaper than Clef but still costs about 1.6x Jev per decision.

Can I use Clef for free on Workers AI?

Yes, up to a limit. Both Workers Free and Workers Paid include 10,000 neurons per day, which covers roughly 1,000 Clef decisions at 455 tokens each. On the Free plan, requests fail once you pass the limit. Workers Paid bills overage and has a $5 monthly minimum.

Is Clef a drop-in replacement for Jev?

Not quite. Clef wraps its response in a result object, which your client must unwrap. Workers AI also returns HTTP 400 when every option key in a choice question contains a slash. Dotted keys fix the slash problem, with the real path kept in the description.

Can I self-host Cloudflare Clef?

Yes. The weights are Apache 2.0 on Hugging Face as Cloudflare/clef and Cloudflare/clef-flash. Clef sits on a 27B base model, so it needs serious GPU memory. Clef-flash uses a 9B base, which asks for much less hardware. Neither was tested self-hosted in this benchmark.


Share this post on:

Send a Webmention

Written about this post on your own site? Send a webmention and it'll show up above once verified.


Previous Post
Radicale vs Baikal: CalDAV That Syncs
Next Post
GitOps Rollback: Argo CD vs Flux

Discussion

Powered by Garrul . Sign in with GitHub or Google, or post anonymously.

Related Posts