Independent experiment · 300 requests · 8 arms · rerun 25 September 2026

Speculative Fan-Out

An agent filling in a product configurator makes eight choices in a row, each one narrowing the next. That loop collapses into three calls and gets more accurate doing it. On both models.

First, what a configurator is

A form that sells a product which does not exist until you finish the form

Think of the page where you build a car. Pick the model, and that decides which engines appear. Pick the engine, and that decides which wheels appear. Industrial suppliers sell the same way, and it is how a large share of B2B manufacturing revenue is actually quoted. A valve is not a thing on a shelf. It is eight choices that together make one part number.

The catch is that the form fights you. Fields stay locked until the thing they depend on is set. Options that were there a second ago vanish when you change an earlier choice. Some combinations are only rejected at the very end, when you press submit.

The catalogue in this experiment is invented, but its shape copies real ones. Eight fields:

1 family
6
valve, coupling, sensor
2 series
61
PC170, SV102
3 body
7
brass, stainless
4 port
12
G20, D80
5 seal
6
NBR, Viton, EPDM
6 actuation
9
24V DC, mains, manual
7 connection
8
flange, NPT, tri-clamp
8 option pack
9
P0 to P8

6 × 61 × 7 × 12 × 6 × 9 × 8 × 9  =  119,532,672

119,532,672
combinations if you just multiply the dropdowns
578,143
that are real, buildable products
0.48%
of the grid is a real product, about 1 in 207

Which is why you cannot just ask once

Almost nothing on that grid is legal, and the form only ever shows you the small slice that still is. Here is a real path through it. The grey squares are what the catalogue holds, the blue ones are what the form actually offers you at that moment, after everything you already picked.

Pick PC170 at step two and the body list drops from seven to three, because that series is only made in brass, polypropylene and stainless. Pick stainless and the seal list drops to two. Each answer rewrites the next question, which is exactly why the normal approach is eight round trips.

The idea

The normal way to drive a configurator is a loop. Ask, set the field, see what changed, ask again. Fan-out looks impossible because the questions are not independent: you cannot ask "which seal" before you know the body.

Unless you ask about hypotheticals. Every question then stands alone, so they all fit in one call, and code walks the tree afterwards using answers it already holds. Branches nobody takes are wasted, but a wasted branch is a few dozen tokens.

SEQUENTIAL                          SPECULATIVE

call 1  which family?               call 1  which family?
        -> PC                               IF family=PC, which series?
call 2  which series?                       IF family=FS, which series?
        -> PC170                            IF body=SS, which seal?
call 3  which body?                         IF body=BR, which seal?
        -> SS                               IF series=PC170, which port?
call 4  which seal?                         ... 29 questions, one call
        -> FF
call 5  which port?                 call 2  the 8 verify questions
        ...                                 "what did the customer name?"
7 calls  4.60s  47.0% right        3 calls  1.89s  55.6% right
7 → 3
model calls per configuration
+8.5 pts
accuracy gained on Jev, p=0.0003
+10.4 pts
accuracy gained on Gemini, p=0.0004
1.89 s
median, against 4.60s for the loop

What it is asked to do

The catalogue above, plus the five things that make a real configurator annoying: fields locked until their dependency is set, options that vanish when an earlier choice changes, one deliberately mislabelled option, a submit that refuses the first press, and one rule that only appears at submit.

Ground truth is built backwards: walk the graph to a valid configuration, then write the customer request from it, so the answer is unambiguous by construction.

"Looking for a process coupling, two inch, Edelstahl, 24 V DC,
 flange, perfluoro, on a steam header."

 family PC  series PC170  body SS  port G20
 seal FF  actuation 024D  connection FLG  pack P0

Five of eight fields stated, three inferred. The answer is PC170-SSFF-G20-024D-FLGP0.

300 requests in four bands: 90 easy, 90 medium, 90 hard, and 30 impossible ones the catalogue forbids, where the only right answer is to refuse. The headline metric is consistent: every field the customer explicitly named is satisfied. An agent that walks the graph legally but picks at random scores 97.4% valid and 0.7% consistent, which is why validity proves nothing.

Result 1

Faster and more accurate at the same time

Requests where every stated field was satisfied, out of 270 solvable. Both models run both architectures, so the comparison isolates the architecture rather than the model.
jev-1.13.0 gemini-3.5-flash-lite random walk, 0.7%
025%50%75%

A8 is the wide speculative call on Gemini. It returned an error on all 300 requests. Section 4 explains why.

ArmConsistentRefused
impossible
Callsp50$ / correct
A1 sequential gem52.2%20%7.33.85s0.00206
A2 sequential jev47.0%60%6.84.60s0.00033
A3 speculative jev55.6%10%3.01.89s0.00048
A4 speculative wide jev55.6%7%1.92.21s0.00456
A5 speculative gem62.6%17%3.03.04s0.00584
A6 spec + verify jev54.8%90%3.01.92s0.00078
A7 spec + verify gem60.7%90%3.03.16s0.00838
A8 spec wide gemfailed 300/300n/an/an/an/a
Paired, same 270 requestsSequentialSpeculativeDiff95% CIp
Jev47.0%55.6%+8.5[+4.1, +13.0]0.0003
Gemini52.2%62.6%+10.4[+5.2, +15.9]0.0004
Jev, both verified47.0%54.8%+7.8[+3.3, +12.2]0.0008
Gemini, both verified52.2%60.7%+8.5[+2.2, +14.4]0.0095

The usual expectation is that batching costs accuracy. It does not here, because each question is still answered against the correctly narrowed option list, just under a hypothesis instead of a commitment.

Result 2

Speculation cannot notice an impossible request

req_0000, impossible by construction

"We need a series PC130, DN80, 1.4308, mains, tri-clamp, Viton,
 stainless fasteners, medium fuel."

tri-clamp is not offered at a flanged size. No valid part exists.

SEQUENTIAL   commits to mains -> reaches the last step ->
             the requested pack is missing -> refuses
SPECULATIVE  answers every branch -> walks a consistent path ->
             swaps in pneumatic -> ships a valid product nobody asked for

Speculation converts "impossible" into "silently different". That is worse than failing, because nothing downstream can tell.

What did not work: asking the model whether the request was self-consistent. 10% recall at 15% false positives, no better than guessing, because the constraints live in the catalogue and not in the request text.

What worked: ask what the customer explicitly named, then compare it in code to what the walk delivered. Those eight questions depend on nothing, so they ride in the same call at no extra round trip.

Share of the 30 impossible requests refused rather than answered with an invented part number. Whiskers are exact 95% intervals, wide because n is 30.
025%50%75%100%

Accuracy cost of adding the verify pass: none measurable. Jev −0.7 points (p=0.77), Gemini −1.9 (p=0.54).

Result 3

Jev is four times better at knowing when it does not know

Verification also refuses about a third of solvable requests, which sounds alarming until you ask what it refused.

Of the solvable requests it refusedJevGemini
refused89 / 27092 / 270
the unverified arm was wrong anyway83 (93%)72 (78%)
the unverified arm was right620
trade ratio14 : 13.6 : 1
net change in correct answers−2−5

Jev gives up 6 correct answers to withdraw 83 wrong part numbers. Gemini gives up 20 to withdraw 72. In a quoting workflow that trade is not close: a wrong part number reaches a customer, a refusal reaches a human.

Result 4

Gemini cannot run the wide version at all

A4 branches over the series too, asking 296 questions instead of 29. On Jev it works. On Gemini it returned 400 INVALID_ARGUMENT on all 300 requests, in about a tenth of a second each, before any generation happens.

This is not the context window. The wide call is roughly 55,000 input tokens against a million-token window. It is the structured-output schema. To guarantee output that matches your schema, Gemini compiles it into a grammar and constrains decoding to it. The grammar is built before generation starts, and every option in every question adds to it.

Where the schema stops compiling, bisected on two question sets. Both die near 700 total options, despite one having two and a half times the questions and half the text.
050010001500

Total options across the whole request. Questions are nearly free; options are what cost.

Question setCeilingOptions thereSchema text
the configurator's real questions126 (127 fails)7305,718 ch
trivial two-option questions329 (330 fails)6582,851 ch
the full wide set, answered by Jev2951,63315,003 ch

Jev has a ceiling too, a different one. Its envelope is 64k tokens for state plus questions. 419 fully labelled questions is about 147,000 characters and returns max_tokens_exceeded. In a separate email test it answered 801 short questions in one call.

And fanning out on Jev is close to free. Same email, one question against 800 in the same call: 0.65s against 1.33s, confidence moving about 0.01, and the main answer unchanged on 39 of 40 emails. The one that flipped had said 0.40 to begin with.

Both models have a ceiling. They are different ceilings. Gemini's is the shape of the answer form, and it hits first at roughly 700 options. Jev's is total text, and it took 1,633 options on the same job.

Result 5

Speculating over everything is wasted money

ArmQuestionsConsistentTotal cost, 300 requests
A3 targeted jev2955.6%$0.07
A4 wide jev29655.6%$0.69
A8 wide gem296failsn/a

Identical accuracy, ten times the questions, nine times the cost. Branch over each step's real dependencies, not over the whole path. The chain is eight steps long but no step depends on all of its predecessors: seal depends only on body, port and actuation only on series. That is what holds the count at 29 instead of thousands, and it is the practical guidance from this work.

Honestly

What this does not show

  • One synthetic catalogue. The generator is published. Nothing here is a claim about any real configurator.
  • A hard precondition. Speculation assumes the option graph can be read once and cached. A configurator that computes options server side on every selection defeats it completely.
  • The verify mechanism was designed after seeing the failure it fixes, so its 90% is in-sample. The interval, 73% to 98%, is what should be quoted.
  • Absolute accuracy is low across the board. On the hard band only three of eight fields are stated. The comparison between arms is the result, not the level.
  • Gemini is the more accurate model here, by about seven points. Jev's case is cost, refusal quality and the absence of a schema ceiling.
  • Arms are only comparable within a single run. No number here is quoted across runs.
  • The error message. Google's docs attribute this failure class to schema complexity, with the wording "too many states for serving". Our logs contain only the bare "Request contains an invalid argument", so that phrase is not quoted as observed.

In one sentence

Batching a dependent decision chain made the agent more accurate, not less.

It held on two independent models, at a third of the calls and half the wall clock, and the one thing it breaks is fixed by eight more questions in the same call. The model-specific part is narrower: Jev withdraws fourteen wrong answers for every correct one it gives up, against Gemini's under four, costs eleven times less per correct configuration, and has no schema ceiling to run into.

Rerun in full on 25 September 2026, all eight arms in one process, and rescored by a second independent implementation that agrees with the original scorer to the decimal.