CONSOLIDATED EVALUATION / SEPTEMBER 2026
Jev / CLM / Kev.
One comparison.
Four evaluation groups, three decision models, and Jev as the reference. Inspect the measured values and the conditions behind them.
BEST Best observed value. Ties included.
Jev: September 26. CLM: September 24. Kev: September 26.
[ 00 / GROUPS ]
One evaluation group per page.
Route the request.
40 tickets, then an exact replay. Compare both CLM head placements.
Choose one next step.
369 judgments across five tasks. Compare quality, errors, consistency, and timing.
Pick the function.
1,253 public cases in three candidate representations. Separate the two official splits.
Keep the agent alive.
Five 60-second courses, with and without a deterministic safety shield.
[ 01 / AT A GLANCE ]
The primary results, together.
Quality and survival
Loading frozen comparisons...
These rows have different units and test different capabilities. They are not combined into an overall model score. The ticket row uses the original CLM CPU-head configuration; its GPU encoder remained active.
[ 02 / READ THE COMPARISON ]
A baseline, not a universal winner.
Jev is the reference in every table. Switch to VS JEV to read quality gaps, latency ratios, and count differences. A latency ratio of 0.25x means one quarter of Jev's measured client latency.
The dashboard consolidates completed runs. CLM was measured on September 24; Kev on September 26. The Jev selector exposes both hosted runs rather than silently merging them. Fixed narrow-decision and BFCL request hashes match across all four runs.
BEST highlights the observed maximum or minimum for that row, including ties. It is not a test of statistical significance. Local servers, hosted network paths, caches, and client location affect response times. The method page explains the conditions and limitations.