One number is doing all the work in coverage of this release: 98.2 percent. It is a compression figure attached to a 27-billion-parameter model whose weights are stored in ternary form — three values instead of the usual large set — and it is presented as the reason the model is interesting.
The figure is probably accurate. The problem is that it is incomplete in a way that makes it unusable, because a percentage is a comparison and no baseline is stated. Ninety-eight point two percent of what, measured how, against which reference.
In this batch of latest LLM models this entry is the clearest case of a claim that reads as a fact and functions as a slogan, and it is worth pulling apart because the same shape appears in every compression story.
On OrcaRouter every latest llm models entry carries a rate you can act on, and the contrast between a measurable rate and an unbaselined percentage is the point of this article.
Why A Compression Percentage Is Meaningless Without Three Numbers
A ratio needs a numerator, a denominator and a definition, and compression claims routinely supply only the first.
The denominator is the reference: the same model at full precision, a comparable dense model at the same parameter count, an earlier version of the same model, or the model’s size as shipped in a particular container format. These produce wildly different percentages, all legitimate.
The definition is what is being measured: the on-disk size of a packaged file, the memory footprint at inference, the theoretical bits per weight, or the compute required per token.
A file’s size on disk includes format overhead; the memory footprint depends on the runtime; the theoretical figure depends on how the packing is done. The same model can honestly be described as 98.2 percent smaller on one measure and half the size on another.
The numerator is usually the only part that gets published, because it is the impressive one and because a bare percentage is easier to put in a headline than a sentence with a reference in it.
What Ternary Weights Actually Buy
The underlying idea is worth understanding rather than dismissing, because it is a real technique with a real payoff.
A conventional model stores each weight as a floating-point number with many possible values. A ternary model constrains each weight to one of three states, which can be encoded in a very small number of bits. The model therefore occupies far less memory and can be moved around far more cheaply, and on hardware that handles the encoding natively, the arithmetic can be faster.
What it does not do is reduce the amount of thinking the model performs. The parameter count is unchanged; the architecture is unchanged; what changed is how the numbers are stored.
A ternary 27-billion-parameter model still routes every token through the same network as its full-precision sibling, so the compute per token is not reduced in proportion to the size reduction. Anyone reading a 98.2 percent figure as “98.2 percent cheaper to run” has made exactly that error.
There is also a quality question, and it is the one that matters for a buyer. Aggressive quantisation degrades some capabilities more than others, and the degradation is not uniform across tasks.
The way to find out how much was lost is not to read a compression percentage but to run the quantised model and the reference model on the same tasks and compare. Which brings the discussion back to the missing baseline: without one, you cannot even identify which comparison to run.
What The Vendor Has And Has Not Said
One further detail is worth stating precisely, because it circulates as fact. The training method behind the model’s quantisation is not something the vendor has spelled out in the material we could read; the description of it as quantisation-aware training, where the model is trained with the reduced precision in the loop rather than compressed afterwards, is an inference drawn from how such models are usually built. It may well be right. It is not a vendor statement, and it should not be repeated as one.
The same discipline applies to the size figures. Different published distributions of a model can differ substantially in on-disk size depending on the packaging format used, so two people quoting two different sizes for “the same model” can both be correct while describing different files. When a size is quoted, the format belongs in the same sentence.

What To Ask Before Believing Any Compression Figure
Four questions convert the claim into something usable, and they take about ten minutes.
Against what reference? If the answer is a specific named model at a specific precision, the percentage is meaningful. If the answer is vague, the number is marketing.
Measured how? File size, resident memory and bits per weight are different quantities. Ask which one, and note that only the third is a property of the model rather than of the container.
At what quality cost? A compression ratio without an evaluation is half a result. The useful form pairs the ratio with a measured difference on a task set that resembles yours.
On what hardware? Ternary storage pays off most where the runtime handles the encoding natively and least where it must be unpacked into a conventional format before use. The same model can be a large win on one deployment and a small one on another. That is not a flaw in the technique; it is a reason the technique is not a universal saving.
Where the Model Fits Regardless
None of this makes the release uninteresting. A capable 27-billion-parameter model in a form that fits in far less memory is genuinely useful, particularly for local deployment, for edge inference, and for any situation where the constraint is memory rather than compute. The category is real and this entry belongs to it.
What it is not is a cheaper way to run a large model in the cloud. If your workload is API calls, the size of the weights on disk is not a line in your bill; what you pay is the provider’s rate, whatever format is underneath. That is why the honest summary of this release is comparative rather than absolute: check it against the full-precision version on your own tasks, and check whether the deployment you actually use can exploit the compressed form at all. If it can, the saving is memory. If the saving you wanted was money, look at the rate.

Sourcing note: Ternary Bonsai 2 27B’s appearance in a third-party public model directory on 18 September 2026, its 27-billion-parameter scale and its ternary weight storage were read from that directory and from the vendor’s own public model organisation page on 28 September 2026. The 98.2 percent figure is the vendor’s own compression claim and is reproduced here as their claim; no baseline, measurement definition or hardware configuration accompanies it in the source material we could read, and no independent evaluation of the compressed model’s quality is quoted because none is published.
The description of the training method as quantisation-aware training is our inference from how comparable models are produced and is explicitly not a vendor statement. The statement that this model is not in OrcaRouter’s catalogue is a statement about our own catalogue on 28 September 2026; no rate of ours is quoted for it, and no per-token price is quoted at all. The discussion of ratio baselines, container formats and hardware dependence is editorial framing rather than measured results. This article names no competing platform and describes none.





