Bonsai 27B: How 1-bit Quantization Put a 27B Multimodal Model on the iPhone 17 Pro
Three numbers tell this story: 54 GB. 18 GB. 3.9 GB. PrismML announced Bonsai 27B on July 12-13: two quantized variants of Qwen3.6 27B, one 1-bit and one ternary. The full-precision 16-bit build of a 27B model needs ~54 GB of memory. Even a 4-bit build lands at 18 GB. Neither fits a phone, and most laptops struggle. PrismML’s 1-bit variant compresses the footprint to 3.9 GB; the ternary variant to 5.9 GB. Both land inside the memory budget of consumer hardware. ...