Independent experiment · 300 requests · 8 arms · rerun 25 September 2026
An agent filling in a product configurator makes eight choices in a row, each one narrowing the next. That loop collapses into three calls and gets more accurate doing it. On both models.
First, what a configurator is
Think of the page where you build a car. Pick the model, and that decides which engines appear. Pick the engine, and that decides which wheels appear. Industrial suppliers sell the same way, and it is how a large share of B2B manufacturing revenue is actually quoted. A valve is not a thing on a shelf. It is eight choices that together make one part number.
The catch is that the form fights you. Fields stay locked until the thing they depend on is set. Options that were there a second ago vanish when you change an earlier choice. Some combinations are only rejected at the very end, when you press submit.
The catalogue in this experiment is invented, but its shape copies real ones. Eight fields:
6 × 61 × 7 × 12 × 6 × 9 × 8 × 9 = 119,532,672
Almost nothing on that grid is legal, and the form only ever shows you the small slice that still is. Here is a real path through it. The grey squares are what the catalogue holds, the blue ones are what the form actually offers you at that moment, after everything you already picked.
Pick PC170 at step two and the body list drops from seven to three, because that series is only made in brass, polypropylene and stainless. Pick stainless and the seal list drops to two. Each answer rewrites the next question, which is exactly why the normal approach is eight round trips.
The idea
The normal way to drive a configurator is a loop. Ask, set the field, see what changed, ask again. Fan-out looks impossible because the questions are not independent: you cannot ask "which seal" before you know the body.
Unless you ask about hypotheticals. Every question then stands alone, so they all fit in one call, and code walks the tree afterwards using answers it already holds. Branches nobody takes are wasted, but a wasted branch is a few dozen tokens.
SEQUENTIAL SPECULATIVE
call 1 which family? call 1 which family?
-> PC IF family=PC, which series?
call 2 which series? IF family=FS, which series?
-> PC170 IF body=SS, which seal?
call 3 which body? IF body=BR, which seal?
-> SS IF series=PC170, which port?
call 4 which seal? ... 29 questions, one call
-> FF
call 5 which port? call 2 the 8 verify questions
... "what did the customer name?"
7 calls 4.60s 47.0% right 3 calls 1.89s 55.6% right
What it is asked to do
The catalogue above, plus the five things that make a real configurator annoying: fields locked until their dependency is set, options that vanish when an earlier choice changes, one deliberately mislabelled option, a submit that refuses the first press, and one rule that only appears at submit.
Ground truth is built backwards: walk the graph to a valid configuration, then write the customer request from it, so the answer is unambiguous by construction.
"Looking for a process coupling, two inch, Edelstahl, 24 V DC, flange, perfluoro, on a steam header." family PC series PC170 body SS port G20 seal FF actuation 024D connection FLG pack P0
Five of eight fields stated, three inferred. The answer is PC170-SSFF-G20-024D-FLGP0.
300 requests in four bands: 90 easy, 90 medium, 90 hard, and 30 impossible ones the catalogue forbids, where the only right answer is to refuse. The headline metric is consistent: every field the customer explicitly named is satisfied. An agent that walks the graph legally but picks at random scores 97.4% valid and 0.7% consistent, which is why validity proves nothing.
Result 1
A8 is the wide speculative call on Gemini. It returned an error on all 300 requests. Section 4 explains why.
| Arm | Consistent | Refused impossible | Calls | p50 | $ / correct |
|---|---|---|---|---|---|
| A1 sequential gem | 52.2% | 20% | 7.3 | 3.85s | 0.00206 |
| A2 sequential jev | 47.0% | 60% | 6.8 | 4.60s | 0.00033 |
| A3 speculative jev | 55.6% | 10% | 3.0 | 1.89s | 0.00048 |
| A4 speculative wide jev | 55.6% | 7% | 1.9 | 2.21s | 0.00456 |
| A5 speculative gem | 62.6% | 17% | 3.0 | 3.04s | 0.00584 |
| A6 spec + verify jev | 54.8% | 90% | 3.0 | 1.92s | 0.00078 |
| A7 spec + verify gem | 60.7% | 90% | 3.0 | 3.16s | 0.00838 |
| A8 spec wide gem | failed 300/300 | n/a | n/a | n/a | n/a |
| Paired, same 270 requests | Sequential | Speculative | Diff | 95% CI | p |
|---|---|---|---|---|---|
| Jev | 47.0% | 55.6% | +8.5 | [+4.1, +13.0] | 0.0003 |
| Gemini | 52.2% | 62.6% | +10.4 | [+5.2, +15.9] | 0.0004 |
| Jev, both verified | 47.0% | 54.8% | +7.8 | [+3.3, +12.2] | 0.0008 |
| Gemini, both verified | 52.2% | 60.7% | +8.5 | [+2.2, +14.4] | 0.0095 |
The usual expectation is that batching costs accuracy. It does not here, because each question is still answered against the correctly narrowed option list, just under a hypothesis instead of a commitment.
Result 2
req_0000, impossible by construction
"We need a series PC130, DN80, 1.4308, mains, tri-clamp, Viton,
stainless fasteners, medium fuel."
tri-clamp is not offered at a flanged size. No valid part exists.
SEQUENTIAL commits to mains -> reaches the last step ->
the requested pack is missing -> refuses
SPECULATIVE answers every branch -> walks a consistent path ->
swaps in pneumatic -> ships a valid product nobody asked for
Speculation converts "impossible" into "silently different". That is worse than failing, because nothing downstream can tell.
What did not work: asking the model whether the request was self-consistent. 10% recall at 15% false positives, no better than guessing, because the constraints live in the catalogue and not in the request text.
What worked: ask what the customer explicitly named, then compare it in code to what the walk delivered. Those eight questions depend on nothing, so they ride in the same call at no extra round trip.
Accuracy cost of adding the verify pass: none measurable. Jev −0.7 points (p=0.77), Gemini −1.9 (p=0.54).
Result 3
Verification also refuses about a third of solvable requests, which sounds alarming until you ask what it refused.
| Of the solvable requests it refused | Jev | Gemini |
|---|---|---|
| refused | 89 / 270 | 92 / 270 |
| the unverified arm was wrong anyway | 83 (93%) | 72 (78%) |
| the unverified arm was right | 6 | 20 |
| trade ratio | 14 : 1 | 3.6 : 1 |
| net change in correct answers | −2 | −5 |
Jev gives up 6 correct answers to withdraw 83 wrong part numbers. Gemini gives up 20 to withdraw 72. In a quoting workflow that trade is not close: a wrong part number reaches a customer, a refusal reaches a human.
Result 4
A4 branches over the series too, asking 296 questions instead of 29. On Jev it works. On Gemini it returned 400 INVALID_ARGUMENT on all 300 requests, in about a tenth of a second each, before any generation happens.
This is not the context window. The wide call is roughly 55,000 input tokens against a million-token window. It is the structured-output schema. To guarantee output that matches your schema, Gemini compiles it into a grammar and constrains decoding to it. The grammar is built before generation starts, and every option in every question adds to it.
Total options across the whole request. Questions are nearly free; options are what cost.
| Question set | Ceiling | Options there | Schema text |
|---|---|---|---|
| the configurator's real questions | 126 (127 fails) | 730 | 5,718 ch |
| trivial two-option questions | 329 (330 fails) | 658 | 2,851 ch |
| the full wide set, answered by Jev | 295 | 1,633 | 15,003 ch |
Jev has a ceiling too, a different one. Its envelope is 64k tokens for state plus questions. 419 fully labelled questions is about 147,000 characters and returns max_tokens_exceeded. In a separate email test it answered 801 short questions in one call.
And fanning out on Jev is close to free. Same email, one question against 800 in the same call: 0.65s against 1.33s, confidence moving about 0.01, and the main answer unchanged on 39 of 40 emails. The one that flipped had said 0.40 to begin with.
Both models have a ceiling. They are different ceilings. Gemini's is the shape of the answer form, and it hits first at roughly 700 options. Jev's is total text, and it took 1,633 options on the same job.
Result 5
| Arm | Questions | Consistent | Total cost, 300 requests |
|---|---|---|---|
| A3 targeted jev | 29 | 55.6% | $0.07 |
| A4 wide jev | 296 | 55.6% | $0.69 |
| A8 wide gem | 296 | fails | n/a |
Identical accuracy, ten times the questions, nine times the cost. Branch over each step's real dependencies, not over the whole path. The chain is eight steps long but no step depends on all of its predecessors: seal depends only on body, port and actuation only on series. That is what holds the count at 29 instead of thousands, and it is the practical guidance from this work.
Honestly
In one sentence
It held on two independent models, at a third of the calls and half the wall clock, and the one thing it breaks is fixed by eight more questions in the same call. The model-specific part is narrower: Jev withdraws fourteen wrong answers for every correct one it gives up, against Gemini's under four, costs eleven times less per correct configuration, and has no schema ceiling to run into.
Rerun in full on 25 September 2026, all eight arms in one process, and rescored by a second independent implementation that agrees with the original scorer to the decimal.