Independent benchmark · October 2026
Two weeks ago we tested TypeSafe's Jev on 1,565 business emails. This week Perplexity and Cloudflare shipped their own decision models with the same API. We ran all four on the exact same emails.
The short version
Background
A chat model writes text. A decision model reads text and returns a probability for every answer you allow. You give it the content and the options. It gives back numbers you can put a threshold on, in one call.
One email in, ten probabilities out
An illustrative example of the task every model got.
Illustrative numbers · the real run used 10 categories and 1,565 emails
That last column is the whole product. If the number is high you let the machine route the email. If it is low you send it to a person. Whether that works depends on one thing: does 0.96 really mean right 96% of the time? That is what this test measures.
Finding 01
All three companies use the same request format. You send some content, a set of named questions, and the options for each. Cloudflare says so directly: you can switch from Jev to Clef "by changing the endpoint and model". We sent the exact same body to all four.
The same request works on all three
Only the address and the model name change.
{
"model": "…",
"state": "Subject: Anfrage DN100 …",
"questions": {
"category": {
"type": "choice",
"instructions": "Classify this email …",
"criteria": { "Order": "…", "Request": "…", … 10 options }
}
}
}
/v1/systemonejev-1.13.0/v1/decisionspplx-decider-v1-27b/ai/run/@cf/cloudflare/…clef · clef-flashSame three question types on every vendor: yes/no, pick one, score on a scale
Perplexity and Cloudflare both published open weights. Both are Qwen language models with a small scoring head on top, run once over the input with no text generation. Perplexity's open-source package is named autojev, serves Jev's endpoint path, and accepts jev-1.13.0 as a model name. Jev's own architecture is not public.
Finding 02
How often each model was right
The second row counts every category equally, so rare ones matter as much as common ones.
The fair way to compare two models on the same emails is to look only where they disagree and count who was right.
| Where they disagreed with Jev | Jev right | Other right | Verdict |
|---|---|---|---|
| Perplexity | 5 | 29 | Perplexity clearly better |
| Clef | 14 | 34 | Clef better |
| Clef-flash | 34 | 23 | No clear difference |
Finding 03 · The main one
All four models are right about 96 to 98% of the time. But look at how often each one is willing to say it is at least 90% sure.
Right vs. sure
Grey: how often the model was right. Color: how often it said it was at least 90% sure.
Clef-flash is right 96% of the time but says it is 90% sure only 3% of the time. Clef is right 98% of the time and sure 43% of the time. Their probabilities are too low. The models are better than they think they are.
Does "80% sure" mean right 80% of the time?
Dashed line: a perfectly honest model. Above the line: the model is underselling itself.
Jev
calibration error 0.025
Perplexity
calibration error 0.031
Clef
calibration error 0.131
Clef-flash
calibration error 0.251
Each dot is a band of answers · bigger dots hold more emails · bands with fewer than 10 emails hidden
Jev and Perplexity sit close to the line: when they say 85%, they mean roughly that. Clef and Clef-flash sit far above it. Clef-flash at "65% sure" was right 99% of the time.
Cloudflare says it trained Clef with a loss designed to improve calibration. Its own public leaderboard lists Jev's calibration score and leaves Clef's blank. On this task, Clef's calibration error is about five times Jev's.
Answers at 99% sure or higher
Out of 1,565. Not one of them was wrong, for any model.
This matters if you want to say "auto-route anything above 99%". With Jev, 759 emails cleared that bar and all 759 were right. Perplexity cleared it 3 times; its highest answer was 99.0%. Clef never went above 97.2% and Clef-flash never above 96.2%. A fixed threshold that works for one model means nothing for another.
Finding 04
Sort each model's answers from most to least confident, then go down the list until mistakes pass 1 in 100. Because this uses each model's own ranking, it works whatever scale its numbers are on.
Share of mail handled automatically at 99% precision
Take the most confident answers first and stop before errors pass 1%.
Perplexity can handle 97% of the inbox at that bar, Clef 96%, Jev 94%. Clef ranks its answers well even though its numbers are too low, so it still does well here. Clef-flash falls to 84%.
How far down the list before the first mistake
The most confident share of mail with zero errors.
Excludes one deliberately hard email that all four models missed.
Perplexity goes through 91% of the mail without a single error, Jev 88%, Clef 80%. Clef-flash makes its first mistake at 30%.
We left one email out of this chart, and only this chart. It is a public tender for office cleaning services that mentions a bill of quantities, so it looks like a tender request but is irrelevant to the business. We wrote it as a deliberately hard trap. All four models fell for it. Perplexity was 97.7% sure, which would end its error-free run at 52% on its own. Because every model got it wrong, removing it treats them all the same, and the chart shows what each model does on the rest of the mail.
Jev's first mistake came at 86% confidence, below all 759 answers it was certain about. That is Jev's remaining advantage: its "certain" is a number you can set a rule on without testing first. Perplexity never goes above 99%, so you have to find its cut-off on your own data.
Finding 05
Time per email
Typical email and the slowest 1 in 100, from the same server.
Measured from one US cloud server · 1,565 calls per model
Cloudflare's announcement says Clef is 2.5x faster than Jev and Clef-flash 13x faster. Their own leaderboard notes that Jev's number includes the internet round trip while Clef's was measured on Cloudflare's own servers. Measured the same way, from the same machine, Clef was 2.2x slower than Jev and Clef-flash slightly slower. Clef's slowest call took 5.8 seconds; Jev and Perplexity never went past 0.7.
Finding 06
Cost per 1,000 emails
From each provider's reported token usage. None of them charge for output.
Perplexity charges $0.04 per million input tokens, Jev about the same, Clef-flash about twice that and Clef six times. For context, two weeks ago the two Gemini models cost $0.80 and $1.79 per 1,000 of the same emails. Perplexity's accuracy (98.0%) sits between theirs (97.5% and 98.5%) at a small fraction of the price.
Finding 07
4 of 1,565
We ran Jev on the same emails on September 17 and again today. Four answers changed. The average confidence moved by less than one percentage point. Cloudflare's leaderboard says Jev's answers "drifted" between two of their runs; on this task it was stable. It also got faster: typical latency fell from 0.60 to 0.29 seconds.
Finding 08
Perplexity and Cloudflare advertise image input. We sent every model a small image and a small PDF, each showing one word, and asked which word it was.
| Model | Image | |
|---|---|---|
| Perplexity | Read it correctly | Rejected |
| Clef | Read it correctly | Rejected |
| Clef-flash | Read it correctly | Rejected |
| Jev | Accepted, but guessed | Accepted, read the raw file as text |
Jev does not support images and does not say so. It accepts the request, treats the encoded image as text, and returns a confident-looking guess: "INVOICE" at about 40% no matter which word was in the picture. Code that sends Jev an image gets a normal answer back and no error. That is worth knowing before you build on it.
Everything else in this report is text only, so all four models saw exactly the same input.
Reference
| Jev | Perplexity | Clef / Clef-flash | |
|---|---|---|---|
| Pin a model version | Yes | No | No |
| Probability precision | 2 decimals | Full | 4 decimals |
| The "confidence" field | Undisclosed formula | Rescaled top probability | Undocumented; matches no published formula |
| Context window | 64k | 262k accepted, trained on 8k | 65k documented; long input silently cut |
| Questions per call | Many | Up to 128 | Up to 64; questions affect each other |
| Rate limit | 1,200 / minute | 10 / second | About 300 / minute |
| Images | No (silent) | Yes | Yes |
| Open weights | No | Yes, 27B | Yes, 27B and 9B |
The practical rules that follow: ignore the vendors' own "confidence" fields and use the probabilities, because each vendor computes confidence differently. Ask one question per call on Clef, because its answers to one question shift with what else is in the request. And log the model name and date on every call, because only Jev lets you pin a version.
Reference
| Claim | What we measured |
|---|---|
| Cloudflare: Clef is 2.5x faster than Jev | 2.2x slower, same machine |
| Cloudflare: Clef-flash is 13x faster than Jev | Slightly slower (0.34s vs 0.29s typical) |
| Cloudflare: trained for calibration | Ranks well, but numbers far too low (error 0.131 and 0.251) |
| Cloudflare leaderboard: Jev drifts | 4 of 1,565 answers changed in 15 days |
| Perplexity: calibrated probabilities | Holds (error 0.031) |
| Perplexity: slightly more accurate than Jev | Holds, 1.6 points ahead |
Honestly
In one sentence
Perplexity's new model is the best all-rounder on this task, Jev is the only one that will tell you it is certain and be right every time, and Cloudflare's Clef is accurate but too modest about it.
Method. 1,565 emails: 1,201 real ones from four company mailboxes and 364 written to cover rare categories, the same set as the September Jev benchmark. Every model got identical text and one "pick one of 10" question with category definitions taken straight from a production prompt. Models: jev-1.13.0, pplx-decider-v1-27b, clef, clef-flash, all called on October 2, 2026 from one US cloud server. Confidence is each model's top probability; the vendors' own confidence fields were ignored. Automation rates use 50 random tie-break orders and were stable. Significance by McNemar's exact test on paired answers. Reference answers for the real emails from Claude Opus 5, which never saw any model's output.