← Nikhil Mudholkar

September 2026 · independent experiment

Speculative fan-out for configurator agents

An agent that fills in an industrial product configurator makes eight choices in a row, and each choice narrows the next. The usual approach is a loop with one model call per choice. I tested asking the what-ifs up front instead: "which body? and IF stainless, which seal? and IF brass, which seal?"

~7 → 3Model calls per configuration.
+8.5 / +10.4Accuracy gain in points for Jev and Gemini. Both p < 0.001.
4.60 → 1.89 sMedian wall clock on Jev.

The problem

A valve is not a thing on a shelf. It is eight choices that together make one part number. In this synthetic catalogue, 119.5 million combinations exist on paper and only 578,143 are real products, about 1 in 207. Each answer changes which options the next field offers. That is why agents use one call per step, and why fan-out looks impossible.

Eight dropdowns, about 120 million combinations, 578,143 valid products

The technique

Ask conditional questions. Each one can be answered without knowing the others, so they all fit in one call. Code then propagates the constraints and walks the tree with answers it already has. Branch only over each step's real dependencies: seal depends on body, not on everything before it. That keeps the tree at 29 questions. Branching over everything (296 questions) gives the same accuracy at nine times the cost.

Sequential calls compared with one fan-out call

What I found

Fewer calls, better answers, on both models. 300 requests, 8 arms. The gain belongs to the architecture, not to the model.

Accuracy for sequential and fan-out on Jev and Gemini

Speculation cannot notice an impossible request. It answers every branch and ships a valid product the customer did not ask for. The fix: in the same call, ask what the customer explicitly named, then compare it in code. Refusal of impossible requests went from 7 to 10% up to 90% at no measurable accuracy cost. Jev withdrew 83 wrong part numbers and lost 6 correct ones. Gemini withdrew 72 and lost 20.

Wrong answers blocked against correct answers lost

Gemini's structured output has a ceiling. The wide call used about 55k of a 1M-token window, and Gemini still rejected all 300 requests with 400 INVALID_ARGUMENT. The limit is the schema: Gemini compiles it into a grammar state machine before generation, and it fails at about 700 total options. Jev answered the whole batch.

Gemini schema limit compared with Jev

The trade. With verification on, Gemini is about 6 points more accurate. Jev costs about 1/11 as much per correct configuration.

Accuracy and cost per correct configuration

Limits

One synthetic catalogue. Speculation needs an option graph that you can read once and cache. A configurator that computes options on the server for each selection defeats it. The verify step was designed after I saw the failure it fixes, so quote its 73 to 98% interval, not the 90% point estimate. Absolute accuracy is low because hard requests state only three of eight fields. The comparison between arms is the result.

On X

The thread

nikhil mudholkar@nikhilmudholkar
𝕏

1/7 I tested Jev and Gemini on 300 customer requests. Asking MORE questions cut the agent from ~7 model calls to 3, and made BOTH models more accurate. But Jev also had a nearly 4x better trade-off when filtering wrong answers. Then Gemini hit a limit that had nothing to do…