September 2026 · independent experiment
Speculative fan-out for configurator agents
An agent that fills in an industrial product configurator makes eight choices in a row, and each choice narrows the next. The usual approach is a loop with one model call per choice. I tested asking the what-ifs up front instead: "which body? and IF stainless, which seal? and IF brass, which seal?"
The problem
A valve is not a thing on a shelf. It is eight choices that together make one part number. In this synthetic catalogue, 119.5 million combinations exist on paper and only 578,143 are real products, about 1 in 207. Each answer changes which options the next field offers. That is why agents use one call per step, and why fan-out looks impossible.

The technique
Ask conditional questions. Each one can be answered without knowing the others, so they all fit in one call. Code then propagates the constraints and walks the tree with answers it already has. Branch only over each step's real dependencies: seal depends on body, not on everything before it. That keeps the tree at 29 questions. Branching over everything (296 questions) gives the same accuracy at nine times the cost.

What I found
Fewer calls, better answers, on both models. 300 requests, 8 arms. The gain belongs to the architecture, not to the model.

Speculation cannot notice an impossible request. It answers every branch and ships a valid product the customer did not ask for. The fix: in the same call, ask what the customer explicitly named, then compare it in code. Refusal of impossible requests went from 7 to 10% up to 90% at no measurable accuracy cost. Jev withdrew 83 wrong part numbers and lost 6 correct ones. Gemini withdrew 72 and lost 20.

Gemini's structured output has a ceiling. The wide call used about 55k of a 1M-token window, and Gemini still rejected all 300 requests with 400 INVALID_ARGUMENT. The limit is the schema: Gemini compiles it into a grammar state machine before generation, and it fails at about 700 total options. Jev answered the whole batch.

The trade. With verification on, Gemini is about 6 points more accurate. Jev costs about 1/11 as much per correct configuration.

Limits
One synthetic catalogue. Speculation needs an option graph that you can read once and cache. A configurator that computes options on the server for each selection defeats it. The verify step was designed after I saw the failure it fixes, so quote its 73 to 98% interval, not the 90% point estimate. Absolute accuracy is low because hard requests state only three of eight fields. The comparison between arms is the result.
On X
The thread
1/7 I tested Jev and Gemini on 300 customer requests. Asking MORE questions cut the agent from ~7 model calls to 3, and made BOTH models more accurate. But Jev also had a nearly 4x better trade-off when filtering wrong answers. Then Gemini hit a limit that had nothing to do…