<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>推理优化 | 杜叔叔网盘</title><description/><link>http://du.kokura.de</link><item><title>JANG: 面向 Apple Silicon 的 MLX 混合精度量化格式• 针对 MoE 与超大模型做按张量分配比特宽度，在接近 MLX 体积下优先保留 attention 与 router 精度，低比特场景下稳定性和效果更强• 支持 Qwen、Nemotron、MiniMax、DeepSeek 等架构，部分模型可在 16 GB 或 64 GB Mac 上运行，甚至实现 397B 级模型在 128 GB Mac 上推理• 原生兼容 MLX 生态，提供推理模式、VLM 支持、bfloat16 自动检测和开发者集成方案，适合在 Apple Silicon 上部署高压缩本地模型</title><link>http://du.kokura.de/posts/341</link><guid isPermaLink="true">http://du.kokura.de/posts/341</guid><pubDate>Wed, 25 Mar 2026 04:51:15 GMT</pubDate><content:encoded>&lt;div class=&quot;tgme_widget_message_forwarded_from accent_color&quot;&gt;Forwarded from &lt;a class=&quot;tgme_widget_message_forwarded_from_name&quot; href=&quot;https://t.me/zhetengsha/5175&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot;&gt;&lt;span&gt;折腾啥&lt;/span&gt;&lt;/a&gt; (&lt;span class=&quot;tgme_widget_message_forwarded_from_author&quot;&gt;小一&lt;/span&gt;)&lt;/div&gt;&lt;u&gt;JANG: 面向 Apple Silicon 的 MLX 混合精度量化格式&lt;/u&gt;&lt;br /&gt;&lt;br /&gt;• 针对 MoE 与超大模型做按张量分配比特宽度，在接近 MLX 体积下优先保留 attention 与 router 精度，低比特场景下稳定性和效果更强&lt;br /&gt;&lt;br /&gt;• 支持 Qwen、Nemotron、MiniMax、DeepSeek 等架构，部分模型可在 16 GB 或 64 GB Mac 上运行，甚至实现 397B 级模型在 128 GB Mac 上推理&lt;br /&gt;&lt;br /&gt;• 原生兼容 MLX 生态，提供推理模式、VLM 支持、bfloat16 自动检测和开发者集成方案，适合在 Apple Silicon 上部署高压缩本地模型&lt;br /&gt;&lt;br /&gt;&lt;a href=&quot;https://github.com/jjang-ai/jangq&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; title=&quot;https://github.com/jjang-ai/jangq&quot;&gt;https://github.com/jjang-ai/jangq&lt;/a&gt;&lt;br /&gt;&lt;br /&gt;&lt;a href=&quot;/search/result?q=%23AppleSilicon&quot; title=&quot;#AppleSilicon&quot;&gt;#AppleSilicon&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23MLX&quot; title=&quot;#MLX&quot;&gt;#MLX&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E6%A8%A1%E5%9E%8B%E9%87%8F%E5%8C%96&quot; title=&quot;#模型量化&quot;&gt;#模型量化&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E6%B7%B7%E5%90%88%E7%B2%BE%E5%BA%A6%E9%87%8F%E5%8C%96&quot; title=&quot;#混合精度量化&quot;&gt;#混合精度量化&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E6%9C%AC%E5%9C%B0%E5%A4%A7%E6%A8%A1%E5%9E%8B&quot; title=&quot;#本地大模型&quot;&gt;#本地大模型&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23MoE&quot; title=&quot;#MoE&quot;&gt;#MoE&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23%E6%8E%A8%E7%90%86%E4%BC%98%E5%8C%96&quot; title=&quot;#推理优化&quot;&gt;#推理优化&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23Mac&quot; title=&quot;#Mac&quot;&gt;#Mac&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23AI&quot; title=&quot;#AI&quot;&gt;#AI&lt;/a&gt; &lt;a href=&quot;/search/result?q=%23GitHub&quot; title=&quot;#GitHub&quot;&gt;#GitHub&lt;/a&gt;&lt;a class=&quot;tgme_widget_message_link_preview&quot; href=&quot;https://github.com/jjang-ai/jangq&quot; target=&quot;_blank&quot; rel=&quot;noopener&quot; title=&quot;JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon - jjang-ai/jangq&quot;&gt;
  
  &lt;div class=&quot;link_preview_site_name accent_color&quot;&gt;GitHub&lt;/div&gt;
  &lt;img class=&quot;link_preview_image&quot; alt=&quot;GitHub - jjang-ai/jangq: JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for…&quot; src=&quot;/static/https://cdn4.telesco.pe/file/W5XUyYm_F7qJZtMeV2KreGLdWrJJPtIQPdmhIog_I2qoA1A9KfUNZajb-D-RL9ZDemj_YNG7Hq78kfiW0CJh8uU3tbpEfKGJdM0SLHJD1NyGLkfz6DCaFOVYbOSeP0-iHute5za49mJ5etnim-UtxeZdHewOxM2tiYOib76GMw2CTXu4BbLLcCrsO33trN2rOB_36qEjfx6yNleYOxGehF52WsosBl8NdP8Vq7lCeJ9qrRP58I1V0oLQeye9guI__YFS2qaroa89ZuX3160dL7QKlSyBy8Ipq1utx-QNqeUtw1VBwiac7qkPOhfaiXE9XLdDO468GJFm1U3R0LVYZw.jpg&quot; width=&quot;1200&quot; height=&quot;630&quot; loading=&quot;eager&quot; /&gt;
  &lt;div class=&quot;link_preview_title&quot;&gt;GitHub - jjang-ai/jangq: JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for…&lt;/div&gt;
  &lt;div class=&quot;link_preview_description&quot;&gt;JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime for Apple Silicon - jjang-ai/jangq&lt;/div&gt;
&lt;/a&gt;</content:encoded></item></channel></rss>