Lab-001 compared a rule-based workflow router against a three-agent hierarchy and reported 40.0% against 23.0%. That comparison changes two things at once: who classifies the query, and whether a second answer attempt is available. So the seventeen point gap cannot be attributed to the topology.
This Lab varies the two factors independently.
| 1 answer attempt | verifier-gated re-answer | |
|---|---|---|
| rule classify | router | router_verified |
| model classify | hier_noverify | hierarchy |
Plus a fifth arm, router_retry: rule classification with one unconditional re-answer and no verifier. That is the arm separating having a second attempt from deciding when to take one.
Everything except the arms is imported from Lab-001 rather than copied: dataset and seed, system prompts, tool registry, verifier prompt, re-answer prompt, rubric, weights, pass threshold and judge. Both original arms are re-run and re-judged in the same session, so every contrast is internal to one run.
What it found
| Arm | Pass | Model calls | Misroute |
|---|---|---|---|
router | 23.0% | 2 | 33% |
router_retry | 29.0% | 4 | 33% |
router_verified | 30.0% | 7.73 | 33% |
hier_noverify | 15.0% | 3 | 68% |
hierarchy | 40.0% | 8.69 | 68% |
Both replicated arms reproduce Lab-001 exactly, 23.0% and 40.0% with 8.69 mean calls for the hierarchy. That is the check that the rebuild is faithful.
A second answer attempt is worth +6.0pp on its own, with no verifier and no extra agent.
Gating that attempt behind a verifier adds +1.0pp, at 7.73 model calls against 4. The verifier rejects the first answer in 95% of runs, so it is an expensive way to almost always retry. Its judgement about when is worth roughly one query in a hundred.
Model classification has no stable sign: -8.0pp alone, +10.0pp combined with verification. The factors interact, so no additive decomposition of the original seventeen points is honest.
The metric was measuring the wrong thing
In both model-classified arms, queries the classifier sent somewhere other than the dataset’s label pass more often than the ones it agreed on, 44.1% against 31.2% in the hierarchy.
The dataset’s category comes from a keyword map built to organise the corpus. The rule-based switch agrees with it by construction, because both are keyword rules. So misroute_rate measures agreement with a label defined for another purpose, not routing quality. Lab-001 reported it as though it were the latter.
Honest limits
- This settles which factor carries the gap in this configuration. It does not rank workflows against multi-agent systems, and every arm’s pass rate is far too low for production use.
- The contrasts are between-arm differences on the same queries rather than terms in a fitted factorial model, so the interaction between the two factors is not separated.
- One deliberate imperfection: the unconditional-retry arm uses Lab-001’s re-answer prompt verbatim, which says “a reviewer flagged the previous answer” when no reviewer did. Rewording it would have confounded the comparison with a prompt difference. That arm gets a mildly misleading instruction, and the trade is recorded rather than hidden.
Reproduce
ollama serve # llama3.2:3b, qwen2.5:14b
python -m labs.lab_006.run # transcripts for all five arms
python -m labs.lab_006.evaluate # judge and score
python -m labs.lab_006.report # re-render RESULTS.md