Liquid AI ships LFM2.5-2.6B, a 2.6B on-device model it says rivals ones nearly four times bigger
Liquid AI released LFM2.5-2.6B, a 2.6-billion-parameter on-device agentic model the company says runs at 220 tokens per second on an Apple M5 Max.
Liquid AI released LFM2.5-2.6B on August 4, 2026, a 2.6-billion-parameter model built to run entirely on-device that the company says rivals models nearly four times its size on tool use and multi-step agentic tasks.
The model — an agentic language model, meaning it is tuned to call tools and carry out multi-step tasks rather than only chat — was pre-trained on roughly 34 trillion tokens and handles a 128,000-token context window extended through a dedicated mid-training phase, Liquid AI said.
On performance, the claims are Liquid AI’s own and are not independently verified. The company says LFM2.5-2.6B runs at 220 tokens per second on an Apple M5 Max and 113 tokens per second on an AMD Ryzen AI Max+ processor, using under 2.5GB of memory, and about 30 tokens per second on a phone. On a single Nvidia H100 GPU it reaches roughly 15,000 output tokens per second at high concurrency, according to the company.
Liquid AI has published both checkpoints on Hugging Face (a base model, LFM2.5-2.6B-Base, and a post-trained version) with day-one support in llama.cpp, MLX, vLLM, SGLang and ONNX.
The pitch fits a broader push to move capable models off cloud servers and onto laptops and phones, where they run without sending data out or racking up per-token API bills. The caveat is that every figure here is a vendor benchmark. Until outside developers reproduce the tool-use and latency numbers on their own hardware, ‘competitive with models nearly four times its size’ stays a claim, not a result.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
