[ J / C / K ] DECISION MODEL LAB

CONSOLIDATED EVALUATION / SEPTEMBER 2026

A bounded decision.
A visible error.

Compare tool choice, trajectory choice, relevance, completion, and file selection on the same 369 fixed requests.

JEV / HOSTED BASELINECLM / PUBLIC 8B HEADKEV / PUBLIC 4B CHECKPOINT
JEV BASELINE
TABLE VALUES
HIGHLIGHT KEY

BEST Best observed value. Ties included.

Jev: September 26. CLM: September 24. Kev: September 26.

VS JEV: accuracy and recall use percentage-point differences; F1 uses F1-point differences; latency uses a ratio to Jev; counts use an absolute difference.

[ 01 / QUALITY ]

Five decisions inside an agent.

Task quality

Loading frozen comparisons...

[ 02 / ERROR COST ]

Keep precision and recall together.

Binary error diagnostics

Loading frozen comparisons...

[ 03 / TIMING ]

What the client waited for.

Median and tail latency

Loading frozen comparisons...

[ 04 / VARIATIONS ]

Consistency is a separate result.

Paired consistency and total correct

Loading frozen comparisons...

[ 05 / OBSERVATIONS ]

A high total can hide missed evidence.

Jev and Kev were perfect on the tool-choice and trajectory-choice fixtures. Kev had the higher positive F1 on relevance and file selection; Jev had the higher subgoal-completion F1. CLM's zero-shot public head had much lower recall in the binary tasks.

There are only 18 relevant files among 225 file judgments. Rejecting all files gives 92% accuracy while retrieving nothing useful. That is why the main table uses positive F1 and recall for binary judgments rather than ranking on the combined correct count.

The 72 paired variants and the file judgments are related observations. The synthetic fixtures were published during the earlier evaluation; training overlap was not audited. These results describe this small fixture set, not unseen production agent traces.