October 2026 · independent benchmark
Four decision models on the same emails
Two weeks after Jev, Perplexity and Cloudflare released decision models with the same request shape. I ran all four on the same 1,565 emails, categories and labels as the Jev vs Gemini benchmark. 6,260 calls, zero errors.
What I found
Perplexity is the strongest model overall. It is the most accurate, the most automatable and the cheapest at $0.062 per 1,000 emails. On true labels alone its lead over Jev is borderline (p = 0.057).

Clef is accurate but badly underconfident. When it says 75%, it is right 98% of the time. It ranks answers well, but you cannot read its numbers as probabilities. Clef-flash is worse and never goes above 0.96.


Jev still has the longest error-free stretch. 759 answers at 99% or higher, all correct. The other models never or almost never return 99%, so a fixed threshold does not transfer between them. Compare them by ranking instead.


Speed and cost. Jev and Perplexity tie on latency (about 0.26 to 0.29 s median). Clef is the slowest and the most expensive at $0.406 per 1,000 emails.

Limits
Claude set the labels for real mail, and the significant gaps weaken on true labels alone. The 364 synthetic emails are cleaner than real mail. No attachments were tested. Perplexity and Clef cannot be version-pinned, so results apply to the hosted models on 2 October 2026.
On X
The thread
1/8 Perplexity and Cloudflare released new decision models yesterday. I ran them against Jev on the same 1,565 business emails from my last benchmark. 6,260 calls. 10 categories. 4 models. The cheapest model was also the most accurate. $0.06 per 1,000 emails🧵