AI & Tech
A 2.6B Model Beats 9B Models at Tool Use in Under 2.5GB of RAM
Liquid AI’s LFM2.5-2.6B runs at 220 tokens per second on an Apple M5 Max and 113 on an AMD Ryzen CPU, in under 2.5GB of memory, with a 131K context and 16 languages. On agentic benchmarks it reports 85.49 IFStruct, 80.07 Multi-IF and 77.83 ToolSandbox — reportedly ahead of Gemma-4-8B and Qwen3.5-9B. Liquid is explicit that this is not a coding or knowledge model. That honesty is the point: it is a dispatcher, and the result suggests tool-calling is a trainable skill largely decoupled from parameter count. The obvious architecture is on-device routing that escalates to a frontier model only when needed.