<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>On-Device Agent on AI Pulse</title>
    <link>https://cuigh.com/tags/on-device-agent/</link>
    <description>Recent content in On-Device Agent on AI Pulse</description>
    <generator>Hugo</generator>
    <language>en-US</language>
    <lastBuildDate>Mon, 13 Jul 2026 09:51:14 +0800</lastBuildDate>
    <atom:link href="https://cuigh.com/tags/on-device-agent/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Bonsai 27B: How 1-bit Quantization Put a 27B Multimodal Model on the iPhone 17 Pro</title>
      <link>https://cuigh.com/posts/bonsai-27b-1bit-on-phone-2026/</link>
      <pubDate>Mon, 13 Jul 2026 09:51:14 +0800</pubDate>
      <guid>https://cuigh.com/posts/bonsai-27b-1bit-on-phone-2026/</guid>
      <description>&lt;p&gt;Three numbers tell this story: 54 GB. 18 GB. 3.9 GB.&lt;/p&gt;
&lt;p&gt;PrismML announced Bonsai 27B on July 12-13: two quantized variants of Qwen3.6 27B, one 1-bit and one ternary. The full-precision 16-bit build of a 27B model needs ~54 GB of memory. Even a 4-bit build lands at 18 GB. Neither fits a phone, and most laptops struggle. PrismML&amp;rsquo;s 1-bit variant compresses the footprint to 3.9 GB; the ternary variant to 5.9 GB. Both land inside the memory budget of consumer hardware.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
