Models

Liquid AI LFM2.5 Beats Rivals in On-Device Benchmarks

A new mobile benchmarking suite from Artificial Analysis and Liquid AI shows that the LFM2.5-2.6B model outperforms larger rivals when tested on actual smartphone hardware.

AlphaSignal19 hrs agoModels
Image: AlphaSignal

Artificial Analysis and Liquid AI have launched a joint benchmarking initiative to evaluate small, 4-bit quantized language models directly on physical smartphones like the iPhone 17 Pro and Galaxy S26 Ultra. Running via llama.cpp, the tests require models to fit within 8 GB of memory. Under a realistic 16K context limit, Liquid AI's LFM2.5-2.6B and Nanbeige4.2-3B tied for the top spot with an average score of 63, outperforming Ornith-1.0-9B at 62 and Qwen3.5 9B Reasoning at 61.

LFM2.5-2.6B demonstrated superior efficiency on an iPhone 17 Pro, answering a 1,024-token prompt in 8.0 seconds using 2.3 GB of memory, compared to Nanbeige4.2-3B's 21.4 seconds and 4.0 GB. The 9B models required over 25 seconds and 6.9 GB. Qwen3.5 9B Reasoning hit the 16K limit on 29% of its generations. With an unrealistic 64K context limit, Ling 3.0 Tiny led with 66, followed by Nanbeige4.2-3B at 65, and both LFM2.5-2.6B and Qwen3.5 9B Reasoning at 64.

Under a 60-second generation cap, LFM2.5-8B-A1B led with 47, followed by LFM2-2.6B-Exp at 45 and Gemma 4 E4B (Non-reasoning) at 44. LFM2.5-2.6B dropped to 38, while Nanbeige4.2-3B fell to 18 due to its slow speed of 14 tokens per second. In specific tasks, Nanbeige4.2-3B scored 76% on BFCL, 96% on MATH-500, and 67% on GPQA Diamond. Qwen3.5 9B Non-reasoning led BFCL at 77% and GPQA Diamond at 79%, while Ornith-1.0-9B hit 15% on AA-Omniscience. LFM2.5-2.6B scored 59% on IFBench, over 90% on MATH-500, and recorded 79% non-hallucination, compared to 33% for Nanbeige4.2-3B and 24% and 1% for the Qwen3.5 9B variants.

Developers can run these tests using Pipette, Liquid AI's open-source benchmarking app. On-device generation times ranged from 0.9 seconds for LFM2.5-230M to 26.7 seconds for Falcon-H1R-7B on a standard prompt. Peak memory at 4K context spanned from 0.4 GB to 6.9 GB for Ornith-1.0-9B and Qwen3.5 9B. Ultimately, these benchmarks show that cloud leaderboards fail to predict how quantized models perform under real-world mobile constraints.

This is our own summary of reporting by AlphaSignal

More in Models