[ J / C / K ] DECISION MODEL LAB

CONSOLIDATED EVALUATION / SEPTEMBER 2026

Jev / CLM / Kev.
One comparison.

Four evaluation groups, three decision models, and Jev as the reference. Inspect the measured values and the conditions behind them.

JEV / HOSTED BASELINECLM / PUBLIC 8B HEADKEV / PUBLIC 4B CHECKPOINT
JEV BASELINE
TABLE VALUES
HIGHLIGHT KEY

BEST Best observed value. Ties included.

Jev: September 26. CLM: September 24. Kev: September 26.

VS JEV: accuracy and recall use percentage-point differences; F1 uses F1-point differences; latency uses a ratio to Jev; counts use an absolute difference.

[ 00 / GROUPS ]

One evaluation group per page.

[ 01 / AT A GLANCE ]

The primary results, together.

Quality and survival

Loading frozen comparisons...

These rows have different units and test different capabilities. They are not combined into an overall model score. The ticket row uses the original CLM CPU-head configuration; its GPU encoder remained active.

[ 02 / READ THE COMPARISON ]

A baseline, not a universal winner.

Jev is the reference in every table. Switch to VS JEV to read quality gaps, latency ratios, and count differences. A latency ratio of 0.25x means one quarter of Jev's measured client latency.

The dashboard consolidates completed runs. CLM was measured on September 24; Kev on September 26. The Jev selector exposes both hosted runs rather than silently merging them. Fixed narrow-decision and BFCL request hashes match across all four runs.

BEST highlights the observed maximum or minimum for that row, including ties. It is not a test of statistical significance. Local servers, hosted network paths, caches, and client location affect response times. The method page explains the conditions and limitations.