Skip to content

Leaderboard · BFCL

Local LLM tool-use leaderboard

The Berkeley Function-Calling Leaderboard (BFCL) tests tool selection and arguments. Our catalog stores Overall Acc, taking the native function-calling (FC) result where available. The exact source snapshot and benchmark version need re-verification; do not compare these legacy scores with the current source board or use them to establish complete agent performance.

Benchmark verification dates are not yet recorded. The model-size refresh does not recheck scores; consult the source leaderboard for current results. Hardware fit below is estimated at Q4_K_M and does not establish the same benchmark performance.

Score is BFCL (recorded overall accuracy on the berkeley function-calling leaderboard.), sourced from gorilla.cs.berkeley.edu/leaderboard. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.

On agentic leaderboards

Tool-call accuracy is one part of an agent's performance. Completing a task also depends on the harness, prompts, available tools and retry budget. Compare whole-task results only when those conditions match; a BFCL score alone cannot tell you whether a local model will replace a hosted coding agent.

Other boards

Or pick by what you need with use-case picks, or browse the full model list.

FAQ

Which tracked model has the highest recorded tool use score?

Kimi K2 Instruct leads our partial catalog at 59.1% on BFCL. This does not establish the current overall leader. It needs ~586.6 GB at Q4_K_M, more than any single tracked device, so it wants a high-memory rig. The hardware figure is an estimate, not a measured benchmark run on that device.

How is this tool use leaderboard scored?

The Berkeley Function-Calling Leaderboard (BFCL) tests tool selection and arguments. Our catalog stores Overall Acc, taking the native function-calling (FC) result where available. The exact source snapshot and benchmark version need re-verification; do not compare these legacy scores with the current source board or use them to establish complete agent performance. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: gorilla.cs.berkeley.edu/leaderboard.

Can I run these models locally?

Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.

Benchmark scores come from gorilla.cs.berkeley.edu/leaderboard and describe its evaluation setup. Memory-data refresh: 2026-09-14; this is not a benchmark verification date. See methodology.