← Nikhil Mudholkar

September 2026 · independent benchmark

Jev vs Gemini on email triage

TypeSafe's Jev is a "System One" decision model. It does not write text. It picks an answer and returns a probability for each option. I ran it against two Google models on 1,565 real and synthetic business emails.

96.4%Jev accuracy. Gemini 3.5 Flash-Lite 97.5%, Gemini 3.8 Flash 98.5%.
$0.08Jev cost per 1,000 emails. Gemini $0.80 and $1.79.
737 / 0Answers at 99%+ confidence, and how many were wrong.

The task

A German B2B supplier receives orders, quote requests, invoice disputes, address changes and auto-replies all day. Each email must go to the right team. All three models got the same 1,565 emails, the same ten categories and the same instructions. About three quarters of the emails are German.

What I found

Jev is a little less accurate. It gets about one more email wrong per hundred than Flash-Lite, and two more than 3.8 Flash. The difference is real but small.

Accuracy of Jev, Gemini 3.5 Flash-Lite and Gemini 3.8 Flash

Jev costs a tenth as much. Cost is measured from actual token use, not list prices.

Cost per 1,000 emails for each model

The slow tail is where it shows. Jev's slowest answer took 1.3 seconds. Flash-Lite's slowest took 57 seconds, and its p99 was 16.2 seconds.

Latency percentiles for each model

The main result: it tells you which answers to trust. Both Gemini models wrote a confidence of 1.0 for most emails and never went below 0.6. Jev's probabilities are calibrated. When it says 85%, it is right 94% of the time. All 737 answers at 99% or higher were correct.

Calibration of Jev's confidence against its accuracy

Sort the answers by confidence and Jev has a perfect record through the most confident 85.5% of the mail. That is a result on this data, not a guaranteed automation rate.

Accuracy as the share of automatically handled mail grows

Limits

Claude Opus 5 set the reference labels for the real emails, so most answers were not checked by a human. On the 364 synthetic emails with known labels, Jev scored 96.2%. No model saw the attachments, and 84% of the real emails had one. This is one task, and it is the task Jev is built for.

On X

The thread

nikhil mudholkar@nikhilmudholkar
𝕏

Jev lost to Gemini on our email classification benchmark. I’m still interested in putting it into production. We tested 1,565 German and English business emails across 10 categories from industrial suppliers. The interesting result wasn’t accuracy. It was where the mistakes…

Diogo Almeida, CEO of TypeSafe, the company that makes Jev, quoted the thread:

Diogo Almeida@CompleteSkeptic
𝕏

yay for private evals and calibrated confidence scores! 🥹 go automation!

nikhil mudholkar@nikhilmudholkar
𝕏

5/8 For email routing, the question is: Which emails can we route automatically, and which need a person? Jev’s most confident 85.5% had zero errors against the reference labels. That’s on this dataset.

17 Sep 2026 · View on X →
17 Sep 2026 · View on X →