Independent benchmark · October 2026

Four decision models, one inbox

Two weeks ago we tested TypeSafe's Jev on 1,565 business emails. This week Perplexity and Cloudflare shipped their own decision models with the same API. We ran all four on the exact same emails.

1,565emails, same as last time
4decision models
6,260calls, zero errors
10categories
Jev (TypeSafe) Perplexity Clef (Cloudflare) Clef-flash (Cloudflare)

The short version

Perplexity wins on most things. Jev is the only one that says "certain".

Background

What a decision model does

A chat model writes text. A decision model reads text and returns a probability for every answer you allow. You give it the content and the options. It gives back numbers you can put a threshold on, in one call.

One email in, ten probabilities out

An illustrative example of the task every model got.

The email"Hello, could you send us a quote for 200 flanged pipe fittings, DN100, delivered to our Hamburg site?"

The questionWhich team should handle this? Pick one of 10.
What comes back
Quote request0.96
Order0.02
Order change0.01
7 others0.01

Illustrative numbers · the real run used 10 categories and 1,565 emails

That last column is the whole product. If the number is high you let the machine route the email. If it is low you send it to a person. Whether that works depends on one thing: does 0.96 really mean right 96% of the time? That is what this test measures.

Finding 01

Three vendors, one API

All three companies use the same request format. You send some content, a set of named questions, and the options for each. Cloudflare says so directly: you can switch from Jev to Clef "by changing the endpoint and model". We sent the exact same body to all four.

The same request works on all three

Only the address and the model name change.

{
  "model": "…",
  "state": "Subject: Anfrage DN100 …",
  "questions": {
    "category": {
      "type": "choice",
      "instructions": "Classify this email …",
      "criteria": { "Order": "…", "Request": "…", … 10 options }
    }
  }
}
TypeSafe
/v1/systemone
jev-1.13.0
Perplexity
/v1/decisions
pplx-decider-v1-27b
Cloudflare
/ai/run/@cf/cloudflare/…
clef · clef-flash

Same three question types on every vendor: yes/no, pick one, score on a scale

Perplexity and Cloudflare both published open weights. Both are Qwen language models with a small scoring head on top, run once over the input with no text generation. Perplexity's open-source package is named autojev, serves Jev's endpoint path, and accepts jev-1.13.0 as a model name. Jev's own architecture is not public.

Finding 02

Perplexity and Clef are more accurate

How often each model was right

The second row counts every category equally, so rare ones matter as much as common ones.

90%92%94%96%98%100%OverallClef-flash 95.7Jev 96.4Clef 97.7Perplexity 98.0Every category counted equallyClef-flash 91.3Jev 91.8Clef 94.3Perplexity 95.2

The fair way to compare two models on the same emails is to look only where they disagree and count who was right.

Where they disagreed with JevJev rightOther rightVerdict
Perplexity529Perplexity clearly better
Clef1434Clef better
Clef-flash3423No clear difference
One caveat matters here. For the 1,201 real emails, the "right answer" came from another AI model, Claude Opus 5, which never saw any model's output. On the 364 emails where we know the answer for certain, Perplexity still leads 11 to 3, but that is just short of statistical significance, and Clef's lead (11 to 8) disappears. So: Perplexity is probably more accurate than Jev, but the labels we fully trust can't quite prove it yet.

Finding 03 · The main one

Three of the four models are never sure

All four models are right about 96 to 98% of the time. But look at how often each one is willing to say it is at least 90% sure.

Right vs. sure

Grey: how often the model was right. Color: how often it said it was at least 90% sure.

0%25%50%75%100%Jevright 96%sure 86%Perplexityright 98%sure 90%Clefright 98%sure 43%Clef-flashright 96%sure 3%

Clef-flash is right 96% of the time but says it is 90% sure only 3% of the time. Clef is right 98% of the time and sure 43% of the time. Their probabilities are too low. The models are better than they think they are.

Does "80% sure" mean right 80% of the time?

Dashed line: a perfectly honest model. Above the line: the model is underselling itself.

Jev

0%0%50%50%100%100%

calibration error 0.025

Perplexity

0%0%50%50%100%100%

calibration error 0.031

Clef

0%0%50%50%100%100%

calibration error 0.131

Clef-flash

0%0%50%50%100%100%

calibration error 0.251

Each dot is a band of answers · bigger dots hold more emails · bands with fewer than 10 emails hidden

Jev and Perplexity sit close to the line: when they say 85%, they mean roughly that. Clef and Clef-flash sit far above it. Clef-flash at "65% sure" was right 99% of the time.

Cloudflare says it trained Clef with a loss designed to improve calibration. Its own public leaderboard lists Jev's calibration score and leaves Clef's blank. On this task, Clef's calibration error is about five times Jev's.

Answers at 99% sure or higher

Out of 1,565. Not one of them was wrong, for any model.

0200400600800Jev759Perplexity3Clef0Clef-flash0

This matters if you want to say "auto-route anything above 99%". With Jev, 759 emails cleared that bar and all 759 were right. Perplexity cleared it 3 times; its highest answer was 99.0%. Clef never went above 97.2% and Clef-flash never above 96.2%. A fixed threshold that works for one model means nothing for another.

Finding 04

So how much can each one automate?

Sort each model's answers from most to least confident, then go down the list until mistakes pass 1 in 100. Because this uses each model's own ranking, it works whatever scale its numbers are on.

Share of mail handled automatically at 99% precision

Take the most confident answers first and stop before errors pass 1%.

0%25%50%75%100%Jev94.4%Perplexity97.0%Clef96.2%Clef-flash83.5%

Perplexity can handle 97% of the inbox at that bar, Clef 96%, Jev 94%. Clef ranks its answers well even though its numbers are too low, so it still does well here. Clef-flash falls to 84%.

How far down the list before the first mistake

The most confident share of mail with zero errors.

95%96%97%98%99%100%0%25%50%75%100%share of mail handled automatically, most confident first →share correctno mistakes through herePerplexity 91.0%Jev 87.9%Clef 80.2%Clef-flash 30.2%

Excludes one deliberately hard email that all four models missed.

Perplexity goes through 91% of the mail without a single error, Jev 88%, Clef 80%. Clef-flash makes its first mistake at 30%.

We left one email out of this chart, and only this chart. It is a public tender for office cleaning services that mentions a bill of quantities, so it looks like a tender request but is irrelevant to the business. We wrote it as a deliberately hard trap. All four models fell for it. Perplexity was 97.7% sure, which would end its error-free run at 52% on its own. Because every model got it wrong, removing it treats them all the same, and the chart shows what each model does on the rest of the mail.

Jev's first mistake came at 86% confidence, below all 759 answers it was certain about. That is Jev's remaining advantage: its "certain" is a number you can set a rule on without testing first. Perplexity never goes above 99%, so you have to find its cut-off on your own data.

Finding 05

Jev and Perplexity tie on speed. Clef is slower than advertised.

Time per email

Typical email and the slowest 1 in 100, from the same server.

0.0s0.4s0.8s1.2s1.6sJevtypical 0.29s1 in 100: 0.47sPerplexitytypical 0.26s1 in 100: 0.45sCleftypical 0.64s1 in 100: 1.56sClef-flashtypical 0.34s1 in 100: 1.16s

Measured from one US cloud server · 1,565 calls per model

Cloudflare's announcement says Clef is 2.5x faster than Jev and Clef-flash 13x faster. Their own leaderboard notes that Jev's number includes the internet round trip while Clef's was measured on Cloudflare's own servers. Measured the same way, from the same machine, Clef was 2.2x slower than Jev and Clef-flash slightly slower. Clef's slowest call took 5.8 seconds; Jev and Perplexity never went past 0.7.

Finding 06

Perplexity is the cheapest

Cost per 1,000 emails

From each provider's reported token usage. None of them charge for output.

$0.00$0.10$0.20$0.30$0.40Jev$0.08Perplexity$0.06Clef$0.41Clef-flash$0.15

Perplexity charges $0.04 per million input tokens, Jev about the same, Clef-flash about twice that and Clef six times. For context, two weeks ago the two Gemini models cost $0.80 and $1.79 per 1,000 of the same emails. Perplexity's accuracy (98.0%) sits between theirs (97.5% and 98.5%) at a small fraction of the price.

Finding 07

Jev barely changed in two weeks

4 of 1,565

We ran Jev on the same emails on September 17 and again today. Four answers changed. The average confidence moved by less than one percentage point. Cloudflare's leaderboard says Jev's answers "drifted" between two of their runs; on this task it was stable. It also got faster: typical latency fell from 0.60 to 0.29 seconds.

Finding 08

Images yes, PDFs no, and one silent failure

Perplexity and Cloudflare advertise image input. We sent every model a small image and a small PDF, each showing one word, and asked which word it was.

ModelImagePDF
PerplexityRead it correctlyRejected
ClefRead it correctlyRejected
Clef-flashRead it correctlyRejected
JevAccepted, but guessedAccepted, read the raw file as text

Jev does not support images and does not say so. It accepts the request, treats the encoded image as text, and returns a confident-looking guess: "INVOICE" at about 40% no matter which word was in the picture. Code that sends Jev an image gets a normal answer back and no error. That is worth knowing before you build on it.

Everything else in this report is text only, so all four models saw exactly the same input.

Reference

API differences that matter

JevPerplexityClef / Clef-flash
Pin a model versionYesNoNo
Probability precision2 decimalsFull4 decimals
The "confidence" fieldUndisclosed formulaRescaled top probabilityUndocumented; matches no published formula
Context window64k262k accepted, trained on 8k65k documented; long input silently cut
Questions per callManyUp to 128Up to 64; questions affect each other
Rate limit1,200 / minute10 / secondAbout 300 / minute
ImagesNo (silent)YesYes
Open weightsNoYes, 27BYes, 27B and 9B

The practical rules that follow: ignore the vendors' own "confidence" fields and use the probabilities, because each vendor computes confidence differently. Ask one question per call on Clef, because its answers to one question shift with what else is in the request. And log the model name and date on every call, because only Jev lets you pin a version.

Reference

The launch claims, checked

ClaimWhat we measured
Cloudflare: Clef is 2.5x faster than Jev2.2x slower, same machine
Cloudflare: Clef-flash is 13x faster than JevSlightly slower (0.34s vs 0.29s typical)
Cloudflare: trained for calibrationRanks well, but numbers far too low (error 0.131 and 0.251)
Cloudflare leaderboard: Jev drifts4 of 1,565 answers changed in 15 days
Perplexity: calibrated probabilitiesHolds (error 0.031)
Perplexity: slightly more accurate than JevHolds, 1.6 points ahead

Honestly

What this doesn't prove

In one sentence

Perplexity's new model is the best all-rounder on this task, Jev is the only one that will tell you it is certain and be right every time, and Cloudflare's Clef is accurate but too modest about it.

Method. 1,565 emails: 1,201 real ones from four company mailboxes and 364 written to cover rare categories, the same set as the September Jev benchmark. Every model got identical text and one "pick one of 10" question with category definitions taken straight from a production prompt. Models: jev-1.13.0, pplx-decider-v1-27b, clef, clef-flash, all called on October 2, 2026 from one US cloud server. Confidence is each model's top probability; the vendors' own confidence fields were ignored. Automation rates use 50 random tie-break orders and were stable. Significance by McNemar's exact test on paired answers. Reference answers for the real emails from Claude Opus 5, which never saw any model's output.