02Agent systems
← The index
When a routing benchmark can measure nothing
Ten pre-registered experiments. Seven negative. Those were the result.
- Made for
- Ten pre-registered experiments
- Years
- 2026
- Status
- Unpublished
I set out to validate a routing design and ended up with a result about the benchmarks instead: one can measure state-conditioned routing only so far as its state is not recoverable from its own transcript. MultiWOZ, SGD and ABCD each fail that. None of ABCD’s 1,004 test dialogues runs more than one workflow. Seven of ten registrations came back negative, including the one that killed my original hypothesis. Not written up yet.