Amartya Gaur
02Agent systems
← The index

When a routing benchmark can measure nothing

Ten pre-registered experiments. Seven negative. Those were the result.

Made for
Ten pre-registered experiments
Years
2026
Status
Unpublished

I set out to validate a routing design and ended up with a result about the benchmarks instead: one can measure state-conditioned routing only so far as its state is not recoverable from its own transcript. MultiWOZ, SGD and ABCD each fail that. None of ABCD’s 1,004 test dialogues runs more than one workflow. Seven of ten registrations came back negative, including the one that killed my original hypothesis. Not written up yet.