CLM / SEPTEMBER 24
Public 8B head.
CLM v0.1-8B over Qwen3-8B embeddings. Projection head and cache on CPU; local GPU encoder. The GPU-head ticket trial is available as a separate condition. No new CLM run was performed for this dashboard.
CONSOLIDATED EVALUATION / SEPTEMBER 2026
Run dates, serving stacks, baseline formulas, source validation, and the limits of a consolidated comparison.
[ 01 / DEPLOYMENTS ]
CLM / SEPTEMBER 24
CLM v0.1-8B over Qwen3-8B embeddings. Projection head and cache on CPU; local GPU encoder. The GPU-head ticket trial is available as a separate condition. No new CLM run was performed for this dashboard.
KEV / SEPTEMBER 26
jaredpalmer/kev-4b at revision 139fdd94f1b6a6ad80cc15e08fcb99cac885a101, code revision 58d94380d4441d2d9fbb7e5e6d30d9a5b93578ae. bf16 serving, fused kernels, CUDA graphs, shipped temperature 2.41, four-state prefix cache. No workload tuning or date preprocessor.
JEV / TWO RECORDED RUNS
Both selected API runs identified themselves as jev-1.13.0. Five narrow-decision predictions changed between runs, and the game outcomes differed. The cause was not established. Select either complete run consistently across the dashboard.
[ 02 / RELATIVE VALUES ]
Accuracy and recall: model percentage minus Jev percentage, in percentage points (pp). A value of -10 pp means ten points below the selected Jev run, not ten percent lower.
F1: model F1 minus Jev F1, in F1 points. Latency: model client latency divided by Jev client latency. Lower ratios are better; 0.25x is one quarter of Jev latency. Counts and probability error: model value minus Jev value. When Jev has zero deaths or errors, a difference remains defined; an artificial ratio is not shown.
Adapter-change rows: each model's accuracy change from its own description view, in pp. VS JEV subtracts Jev's adapter change, also in pp. These signed changes have no BEST highlight.
Jev's relative value is zero for differences and 1.00x for latency ratios. All cells retain the measured absolute value in either mode. Differences in quality are calculated from raw counts where available; the source timing precision is retained.
[ 03 / VALIDATION ]
The builder verifies all 40 source-file checksums. For narrow decisions and each BFCL view, it compares ordered IDs, states, instructions, options, gold labels, and request hashes across CLM, Kev, and both Jev runs. Binary F1, recall, precision, error counts, and Brier are derived from the frozen rows.
Ticket fixture definitions and seeded order are identical. The two newer ticket runs retain full states and a verified suite hash. Older ticket rows retain only their expected labels and answers, so the source-code and label-order checks are weaker evidence than full wire-request identity.
All selected game runs are checked for five seeds, 60 seconds, six requests in flight, original course, labeled prompt, shield setting, survival count, and deaths. The 14-run September 26 ledger must be complete. No additional model calls are needed to build or open this dashboard.
[ 04 / LIMITS ]
The cells pool completed evaluations from two dates, not one simultaneous three-model experiment. Fixed-request identity supports comparing answers on the same tasks, but it does not equalize cache design, client region, software stack, or network transit. Original ticket clients were in Europe; the later decision, BFCL, game, and Kev comparison clients used US-KS-2.
BEST means the best reported value for that metric among the displayed deployments. Exact ties are marked together; no significance test or universal model ranking is implied. Public fixtures may overlap training data. Related variants, imbalanced labels, and five game seeds limit generalization.
CLM's published coding-verifier results use fine-tuned heads and different tasks. The CLM team's BFCL chart and Kev's published out-of-domain datasets use different protocols. Neither is mixed into these consolidated evaluation tables.
The interrupted Jev prose-schema attempt was excluded from final scores and rerun in full after HTTP 520 retry handling was added. Completed fixed suites recorded no retries. One Jev prose probability distribution sums to 0.99 because of two-decimal rounding; the original report retains it without renormalization.
[ 05 / HOSTING ]
The deployment uses Cloudflare Workers Static Assets: HTML, CSS, browser JavaScript, and frozen JSON. It contains no model credential or live model endpoint integration. The evaluation instances were removed after their source reports were verified. No GPU instance is needed or created for this dashboard.