PrismML releases Bonsai 2 27B under Apache 2.0 for local AI
PrismML released downloadable Bonsai 2 27B weights under Apache 2.0, positioning the compact multimodal model for local AI while its current GGUF files require PrismML's llama.cpp fork.
PrismML released Bonsai 2 27B on September 17, making the model weights available to download under Apache 2.0. The company positions the compact multimodal ternary model, derived from Qwen3.8-27B and distributed in GGUF files, for local assistants, coding, vision and agentic workloads.
The model repository lists a 5.95 GB PTQ1_0 language-model file, a 7.21 GB PQ2_0 alternative and an optional 0.63 GB Q8_0 vision pack. Those are PrismML’s published file sizes, not complete system-memory requirements. Practical use also depends on the selected packing, context length, cache, runtime and whether the vision component is loaded.
PrismML’s materials give different parameter totals. The announcement describes Bonsai 2 as a 27.8-billion-parameter model, while the model card lists 27.36 billion total parameters, including a 460-million-parameter vision tower. The company does not explain the difference.
Qwen’s model card for the base model identifies Qwen3.8-27B as a causal language model with a vision encoder and native image and video understanding. For context, see the Qwen3.8-27B base-model release.
The company also reports two benchmark summaries. Its announcement gives Bonsai 2 an aggregate score of 83.9 across 20 benchmarks, compared with 85.4 for Qwen3.8-27B. The Bonsai model card reports 84.78 versus 86.32 across 14 thinking-mode benchmarks. Both produce PrismML’s claimed 98.2% retention, but the company does not explain how the suites relate, and the results have not been independently reproduced.
The performance claims need similar qualification. PrismML says the model reaches up to 143 tokens per second on an RTX 5090, while its documented batch-one TG128 table lists 129.9 tokens per second for PQ2_0 and 120.5 for PTQ1_0. The company also publishes hardware-specific throughput and per-token energy measurements that were not independently tested for the research package.
Running the repository’s PTQ1_0 and PQ2_0 GGUF files through llama.cpp currently requires PrismML’s custom fork. The model card says stock llama.cpp rejects those two file types; it also warns that stock llama.cpp can load Q2_0 without the required runtime transform and produce invalid output.
More news

AWS details Grok 4.6 access through Amazon Bedrock

Apptopia estimates Meta’s Muse led ChatGPT in early iOS downloads

OpenAI explains V7 Go’s source-linked memory for enterprise agents
