CONSOLIDATED EVALUATION / SEPTEMBER 2026
Which function
fits the request?
Public BFCL function-name selection with descriptions, JSON schemas, and prose parameters. One consolidated table for each measure.
BEST Best observed value. Ties included.
Jev: September 26. CLM: September 24. Kev: September 26.
[ 01 / THREE INPUT VIEWS ]
The same functions,
three candidate adapters.
Function-selection accuracy
Loading frozen comparisons...
[ 02 / TIMING ]
Response times by representation.
Client latency for each view
Loading frozen comparisons...
[ 03 / SPLITS ]
Inspect each official split.
[ 04 / SENSITIVITY ]
Measure the adapter effect.
Change within each model
Loading frozen comparisons...
[ 05 / SCOPE ]
Function names, not complete calls.
Each view uses all 200 multiple and 1,053 live_multiple cases, with official gold function names and original candidate order. The dataset is pinned to Gorilla commit f7cf7359b7ac615a0b294831c5ba2bc95ee4a000. No generated arguments, execution, parallel calls, memory, or overall BFCL leaderboard score is evaluated.
Jev leads the three views in both recorded runs. Kev is near 97% with descriptions and full JSON, but falls to 84.52% with prose parameters. CLM has a different ordering: descriptions perform best, then prose, then full JSON. A candidate adapter that helps one model can hurt another.
The description and prose adapters omit some schema fields. Differences therefore combine representation and information content. Prose was a post-hoc addition in the old CLM evaluation, but all three views were specified before the Kev run. The CLM team's published 95.2% chart uses an unavailable exact protocol and is not treated as an equivalent row here.