Skip to content

Leaderboard · BFCL

Best local LLMs for tool use and agents

The Berkeley Function-Calling Leaderboard (BFCL) scores whether a model picks the right tool and calls it with the right arguments, across simple, parallel and multi-call invocations. We use its Overall Acc column, taking the native function-calling (FC) result where a model has one. It is the closest open, sourced measure of how agent-ready a local model is.

Score is BFCL (overall acc on the berkeley function-calling leaderboard.), sourced from gorilla.cs.berkeley.edu/leaderboard. The "runs on" column is the lightest tracked device that runs the model at Q4_K_M; tap a model for the full hardware list. Coverage is partial: only models we verified against the canonical board appear.

On agentic leaderboards

BFCL tests a concrete capability: does the model call the right tool with the right arguments? Newer agent-focused boards such as Agent Arena rank whole-task orchestration, where today's leaders are still frontier cloud models and open-weight models are well behind. BFCL is the open, sourced signal that covers the local models you can actually run.

Other boards

Or pick by what you need with use-case picks, or browse the full model list.

FAQ

What is the best local model for tool use?

Kimi K2 Instruct tops this board at 59.1% on BFCL. It needs ~586.6 GB at Q4_K_M, more than any single tracked device, so it wants a high-memory rig.

How is this tool use leaderboard scored?

The Berkeley Function-Calling Leaderboard (BFCL) scores whether a model picks the right tool and calls it with the right arguments, across simple, parallel and multi-call invocations. We use its Overall Acc column, taking the native function-calling (FC) result where a model has one. It is the closest open, sourced measure of how agent-ready a local model is. We only list a model after verifying its score against the canonical leaderboard for that exact model, so coverage is partial by design. Source: gorilla.cs.berkeley.edu/leaderboard.

Can I run these models locally?

Every row shows the lightest tracked device that runs the model at Q4_K_M and links to the full hardware list, so you see the ranking and your hardware fit together.

Benchmark scores are third-party, sourced from gorilla.cs.berkeley.edu/leaderboard, and reflect the full-precision model, not a specific quant. Memory and fit are computed at Q4_K_M, validated 2026-08-03. See methodology.