PrismML Compresses Qwen3.8 27B to 5.93 GB

PrismML's Ternary Bonsai 2 27B uses ternary weights to shrink Qwen3.8 27B from 53.80 GB to 5.93 GB.

PrismML Compresses Qwen3.8 27B to 5.93 GB

Image: aiweekly.co

PrismML has introduced Ternary Bonsai 2 27B, a compressed version of Alibaba's Qwen3.8 27B language model. According to MarkTechPost, the model reduces the original size from 53.80 GB to 5.93 GB by quantizing every weight to one of three values: -1, 0, or +1. This ternary approach drastically cuts memory requirements while aiming to preserve performance.

The technique, known as ternary quantization, represents weights with just two bits, enabling significant compression. PrismML's implementation is part of a growing trend to make large language models more accessible on consumer hardware. The company has not yet released detailed benchmarks, but the size reduction alone marks a notable engineering achievement.

While the exact performance trade-offs remain unclear, the development highlights ongoing efforts to optimize AI models for edge devices and lower-resource environments. Further details are expected from PrismML or independent evaluations.

❓ Frequently Asked Questions

What is Ternary Bonsai 2 27B?

It is a compressed version of Alibaba's Qwen3.8 27B model developed by PrismML, using ternary weights to reduce its size from 53.80 GB to 5.93 GB.

How does ternary quantization work?

Ternary quantization constrains each weight to one of three values: -1, 0, or +1, which allows each weight to be stored in about two bits, greatly reducing memory usage.

What are the performance implications?

The exact performance impact is not yet publicly detailed, but the compression aims to maintain model quality while making it more deployable on limited hardware.

📰 Source:
aiweekly.co →
Share: