<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: anemll</title><link>https://news.ycombinator.com/user?id=anemll</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 13 Sep 2026 08:06:27 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=anemll" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by anemll in "Getting 50 GB/S Back from the Apple Neural Engine"]]></title><description><![CDATA[
<p>Not all systems affected
M1 and M5MAX are OK<p><a href="https://x.com/anemll/status/2098454204478366132?s=20" rel="nofollow">https://x.com/anemll/status/2098454204478366132?s=20</a></p>
]]></description><pubDate>Sun, 13 Sep 2026 02:08:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49679242</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=49679242</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49679242</guid></item><item><title><![CDATA[New comment by anemll in "TurboQuant KV Compression and SSD Expert Streaming for M5 Pro and IOS"]]></title><description><![CDATA[
<p>Check it out,
you might be able to speed it up using this
<a href="https://github.com/Anemll/anemll-flash-mlx" rel="nofollow">https://github.com/Anemll/anemll-flash-mlx</a>
<a href="https://x.com/anemll/status/2038684375425200360" rel="nofollow">https://x.com/anemll/status/2038684375425200360</a></p>
]]></description><pubDate>Wed, 01 Apr 2026 20:03:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=47605823</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47605823</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47605823</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>17B includes 10 expert plus one shared. So actual size of the expert is much smaller</p>
]]></description><pubDate>Tue, 24 Mar 2026 03:40:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=47498363</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47498363</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47498363</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>Check my repo, I had added some support for GUFF/untloth, Q3,Q5/Q8
<a href="https://github.com/Anemll/flash-moe/blob/iOS-App/docs/gguf-hybrid-bringup-log.md" rel="nofollow">https://github.com/Anemll/flash-moe/blob/iOS-App/docs/gguf-h...</a></p>
]]></description><pubDate>Mon, 23 Mar 2026 19:19:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=47493921</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47493921</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47493921</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>Thanks for posting this, that's how I first found out about Dan's experiment!
SSD speed doubled in the M5P/M generation, that makes it usable!
I think one paper under the radar is "KV Prediction for Improved Time to First Token" <a href="https://arxiv.org/abs/2410.08391" rel="nofollow">https://arxiv.org/abs/2410.08391</a> which hopefully can help with prefill for Flash streaming.</p>
]]></description><pubDate>Mon, 23 Mar 2026 18:48:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=47493564</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47493564</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47493564</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>SSD streaming to compute units is new.
M4 max can do 15 t/s with its 15GB/s drives</p>
]]></description><pubDate>Mon, 23 Mar 2026 18:40:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=47493458</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47493458</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47493458</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>Yes, SSD speed is critical though. The repo has macOS builds for CLI and Desktop.
It's early stages though. M4 Max gets 10-15 TPS on 400B depending on quantization. Compute is an issue too; a lot of code is PoC level.</p>
]]></description><pubDate>Mon, 23 Mar 2026 18:39:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=47493446</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47493446</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47493446</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>multiple NAND, and apple already used it in Mac Studio.
Plus better cooling</p>
]]></description><pubDate>Mon, 23 Mar 2026 17:52:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=47492835</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47492835</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47492835</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>both, tbh</p>
]]></description><pubDate>Mon, 23 Mar 2026 17:51:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=47492824</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47492824</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47492824</guid></item><item><title><![CDATA[New comment by anemll in "iPhone 17 Pro Demonstrated Running a 400B LLM"]]></title><description><![CDATA[
<p>Probably 2x speed for Mac Studio this year if they do double NAND ( or quad?)</p>
]]></description><pubDate>Mon, 23 Mar 2026 17:49:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=47492792</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47492792</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47492792</guid></item><item><title><![CDATA[iPhone 17 Pro Demonstrated Running a 400B LLM]]></title><description><![CDATA[
<p><a href="https://xcancel.com/anemll/status/2035901335984611412" rel="nofollow">https://xcancel.com/anemll/status/2035901335984611412</a></p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47490070">https://news.ycombinator.com/item?id=47490070</a></p>
<p>Points: 713</p>
<p># Comments: 326</p>
]]></description><pubDate>Mon, 23 Mar 2026 14:30:10 +0000</pubDate><link>https://twitter.com/anemll/status/2035901335984611412</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=47490070</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47490070</guid></item><item><title><![CDATA[New comment by anemll in "macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt"]]></title><description><![CDATA[
<p>Tensor Parallel test with RDMA last week
<a href="https://x.com/anemll/status/1996349871260107102" rel="nofollow">https://x.com/anemll/status/1996349871260107102</a><p>Note fast sync workaround</p>
]]></description><pubDate>Sat, 13 Dec 2025 03:51:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=46251822</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=46251822</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46251822</guid></item><item><title><![CDATA[New comment by anemll in "RDMA over Thunderbolt 5 on Apple Silicon – 14µs latency"]]></title><description><![CDATA[
<p>In macOS 26.2 (Tahoe) beta, Apple introduced a low-latency Thunderbolt 5 RDMA driver, enabling up to 80 Gb/s bidirectional bandwidth for Mac clustering—ideal for distributed ML on Apple Silicon. It's optimized for low latency, delivering ~14 Gbps throughput at 4K MTU.
My tests (M4 Pro to M3 Ultra): Stock ibv_uc_pingpong achieved ~14 µs round-trip for 4K packets (requires GID index setup). Custom C++ variant hit 6-13 µs/iter: <a href="https://x.com/anemll/status/1993192776897642942" rel="nofollow">https://x.com/anemll/status/1993192776897642942</a>
Code and details:
<a href="https://github.com/Anemll/mlx-rdma/blob/anemll-rdma/ibv_roundtrip.cpp" rel="nofollow">https://github.com/Anemll/mlx-rdma/blob/anemll-rdma/ibv_roun...</a>
<a href="https://github.com/Anemll/mlx-rdma/blob/anemll-rdma/ibv_roundtrip.md" rel="nofollow">https://github.com/Anemll/mlx-rdma/blob/anemll-rdma/ibv_roun...</a> (includes steps to enable RDMA in macOS Recovery OS terminal)
Theoretically, this accelerates pipeline parallelism (faster layer handoffs) and tensor parallelism (low-overhead sharding) on GPUs, with potential extensions to ANE for real-time AI workflows.</p>
]]></description><pubDate>Tue, 25 Nov 2025 17:24:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=46048148</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=46048148</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46048148</guid></item><item><title><![CDATA[RDMA over Thunderbolt 5 on Apple Silicon – 14µs latency]]></title><description><![CDATA[
<p>Article URL: <a href="https://twitter.com/anemll/status/1993182652204187929">https://twitter.com/anemll/status/1993182652204187929</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46048147">https://news.ycombinator.com/item?id=46048147</a></p>
<p>Points: 6</p>
<p># Comments: 1</p>
]]></description><pubDate>Tue, 25 Nov 2025 17:24:03 +0000</pubDate><link>https://twitter.com/anemll/status/1993182652204187929</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=46048147</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46048147</guid></item><item><title><![CDATA[New comment by anemll in "Qwen 3 now supports ARM and MLX"]]></title><description><![CDATA[
<p>It’s also supported in Apple Neural Engine
<a href="https://github.com/Anemll/Anemll" rel="nofollow">https://github.com/Anemll/Anemll</a></p>
]]></description><pubDate>Sun, 14 Sep 2025 23:55:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=45244563</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=45244563</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45244563</guid></item><item><title><![CDATA[Anemll adds Qwen3 support for Apple neural engine]]></title><description><![CDATA[
<p>Article URL: <a href="https://twitter.com/anemll/status/1935779931822420447">https://twitter.com/anemll/status/1935779931822420447</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44323416">https://news.ycombinator.com/item?id=44323416</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Thu, 19 Jun 2025 23:22:43 +0000</pubDate><link>https://twitter.com/anemll/status/1935779931822420447</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=44323416</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44323416</guid></item><item><title><![CDATA[New comment by anemll in "Run LLMs on Apple Neural Engine (ANE)"]]></title><description><![CDATA[
<p>We can ran 2000 or 4000 context with ANE</p>
]]></description><pubDate>Wed, 07 May 2025 16:15:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=43917545</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=43917545</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43917545</guid></item><item><title><![CDATA[New comment by anemll in "Run LLMs on Apple Neural Engine (ANE)"]]></title><description><![CDATA[
<p>M4 max should work at 120GB for ANE and 500+ for GPU. So GPU will be 3-4 times faster for anything over 1-3B. ANE is likely as fast for prefill due to higher FLOPs</p>
]]></description><pubDate>Wed, 07 May 2025 16:14:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=43917532</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=43917532</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43917532</guid></item><item><title><![CDATA[New comment by anemll in "Run LLMs on Apple Neural Engine (ANE)"]]></title><description><![CDATA[
<p>Right.I was thinking about it, you still need batch refill, however, Apple Core ML tools were failing for attention activations quantization.  Long context, pre-fill is still compute bound.</p>
]]></description><pubDate>Sun, 04 May 2025 14:28:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=43886921</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=43886921</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43886921</guid></item><item><title><![CDATA[New comment by anemll in "Run LLMs on Apple Neural Engine (ANE)"]]></title><description><![CDATA[
<p>Yes for GPU, however ANE only supports FP16 plus integers. M4/A17 added  accelerated int8  that is twice faster than FP16</p>
]]></description><pubDate>Sun, 04 May 2025 14:22:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=43886877</link><dc:creator>anemll</dc:creator><comments>https://news.ycombinator.com/item?id=43886877</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43886877</guid></item></channel></rss>