CONSOLIDATED EVALUATION / SEPTEMBER 2026
Route the ticket.
Then route it again.
Support-ticket routing under new-state and exact-replay conditions. Accuracy, median latency, and tail latency in one table per condition.
BEST Best observed value. Ties included.
Jev: September 26. CLM: September 24. Kev: September 26.
[ 01 / CONDITION ]
Forty tickets. Four teams.
The encoder uses a GPU in both conditions. Only CLM's ticket column changes.
Each model chooses sales, billing, technical support, or account support. The fixture text, option descriptions, and fixed ticket order are retained across the runs.
[ 02 / NEW STATES ]
First encounter with each ticket.
New-state accuracy and latency
Loading frozen comparisons...
[ 03 / REPLAY ]
The same forty tickets again.
Exact-replay accuracy and latency
Loading frozen comparisons...
[ 04 / OBSERVATIONS ]
The cache changes the measurement.
Jev and Kev routed every ticket correctly in both passes. CLM routed 27/40 correctly with its CPU head and 26/40 with its GPU head. Moving the head changed one close billing-versus-sales choice; it was an exploratory run, not a separately tuned checkpoint.
CLM's original replay median was 1.95 ms, reflecting its local vector and answer caches. Kev's default prefix cache holds four states; cycling through forty tickets exceeds that capacity. The replay timings measure different cache conditions.
The original ticket clients ran from a European pod. The September 26 clients ran from US-KS-2. Their hosted Jev latency is therefore especially sensitive to client location. Older ticket results did not retain full request text: fixture-code equality and gold-label order are verified, while the two newer runs also retain and validate every state.