September 2026 · independent benchmark
Jev vs Gemini on email triage
TypeSafe's Jev is a "System One" decision model. It does not write text. It picks an answer and returns a probability for each option. I ran it against two Google models on 1,565 real and synthetic business emails.
The task
A German B2B supplier receives orders, quote requests, invoice disputes, address changes and auto-replies all day. Each email must go to the right team. All three models got the same 1,565 emails, the same ten categories and the same instructions. About three quarters of the emails are German.
What I found
Jev is a little less accurate. It gets about one more email wrong per hundred than Flash-Lite, and two more than 3.8 Flash. The difference is real but small.

Jev costs a tenth as much. Cost is measured from actual token use, not list prices.

The slow tail is where it shows. Jev's slowest answer took 1.3 seconds. Flash-Lite's slowest took 57 seconds, and its p99 was 16.2 seconds.

The main result: it tells you which answers to trust. Both Gemini models wrote a confidence of 1.0 for most emails and never went below 0.6. Jev's probabilities are calibrated. When it says 85%, it is right 94% of the time. All 737 answers at 99% or higher were correct.

Sort the answers by confidence and Jev has a perfect record through the most confident 85.5% of the mail. That is a result on this data, not a guaranteed automation rate.

Limits
Claude Opus 5 set the reference labels for the real emails, so most answers were not checked by a human. On the 364 synthetic emails with known labels, Jev scored 96.2%. No model saw the attachments, and 84% of the real emails had one. This is one task, and it is the task Jev is built for.
On X
The thread
Jev lost to Gemini on our email classification benchmark. I’m still interested in putting it into production. We tested 1,565 German and English business emails across 10 categories from industrial suppliers. The interesting result wasn’t accuracy. It was where the mistakes…
Diogo Almeida, CEO of TypeSafe, the company that makes Jev, quoted the thread:
yay for private evals and calibrated confidence scores! 🥹 go automation!
5/8 For email routing, the question is: Which emails can we route automatically, and which need a person? Jev’s most confident 85.5% had zero errors against the reference labels. That’s on this dataset.