CONSOLIDATED EVALUATION / SEPTEMBER 2026
A bounded decision.
A visible error.
Compare tool choice, trajectory choice, relevance, completion, and file selection on the same 369 fixed requests.
BEST Best observed value. Ties included.
Jev: September 26. CLM: September 24. Kev: September 26.
[ 01 / QUALITY ]
Five decisions inside an agent.
Task quality
Loading frozen comparisons...
[ 02 / ERROR COST ]
Keep precision and recall together.
Binary error diagnostics
Loading frozen comparisons...
[ 03 / TIMING ]
What the client waited for.
Median and tail latency
Loading frozen comparisons...
[ 04 / VARIATIONS ]
Consistency is a separate result.
Paired consistency and total correct
Loading frozen comparisons...
[ 05 / OBSERVATIONS ]
A high total can hide missed evidence.
Jev and Kev were perfect on the tool-choice and trajectory-choice fixtures. Kev had the higher positive F1 on relevance and file selection; Jev had the higher subgoal-completion F1. CLM's zero-shot public head had much lower recall in the binary tasks.
There are only 18 relevant files among 225 file judgments. Rejecting all files gives 92% accuracy while retrieving nothing useful. That is why the main table uses positive F1 and recall for binary judgments rather than ranking on the combined correct count.
The 72 paired variants and the file judgments are related observations. The synthetic fixtures were published during the earlier evaluation; training overlap was not audited. These results describe this small fixture set, not unseen production agent traces.