<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: eiln</title><link>https://news.ycombinator.com/user?id=eiln</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 13 Sep 2026 05:07:26 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=eiln" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by eiln in "Getting 50 GB/S Back from the Apple Neural Engine"]]></title><description><![CDATA[
<p>RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.</p>
]]></description><pubDate>Thu, 10 Sep 2026 00:16:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49636480</link><dc:creator>eiln</dc:creator><comments>https://news.ycombinator.com/item?id=49636480</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49636480</guid></item><item><title><![CDATA[Getting 50 GB/S Back from the Apple Neural Engine]]></title><description><![CDATA[
<p>Article URL: <a href="https://eiln.github.io/posts/ane-dma.html">https://eiln.github.io/posts/ane-dma.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49636479">https://news.ycombinator.com/item?id=49636479</a></p>
<p>Points: 108</p>
<p># Comments: 22</p>
]]></description><pubDate>Thu, 10 Sep 2026 00:16:10 +0000</pubDate><link>https://eiln.github.io/posts/ane-dma.html</link><dc:creator>eiln</dc:creator><comments>https://news.ycombinator.com/item?id=49636479</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49636479</guid></item><item><title><![CDATA[New comment by eiln in "Getting 50 GB/S Back Out of the Neural Engine"]]></title><description><![CDATA[
<p>RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.</p>
]]></description><pubDate>Wed, 09 Sep 2026 03:43:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49620643</link><dc:creator>eiln</dc:creator><comments>https://news.ycombinator.com/item?id=49620643</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49620643</guid></item><item><title><![CDATA[Getting 50 GB/S Back Out of the Neural Engine]]></title><description><![CDATA[
<p>Article URL: <a href="https://eiln.github.io/posts/ane-dma.html">https://eiln.github.io/posts/ane-dma.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49620642">https://news.ycombinator.com/item?id=49620642</a></p>
<p>Points: 3</p>
<p># Comments: 1</p>
]]></description><pubDate>Wed, 09 Sep 2026 03:43:16 +0000</pubDate><link>https://eiln.github.io/posts/ane-dma.html</link><dc:creator>eiln</dc:creator><comments>https://news.ycombinator.com/item?id=49620642</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49620642</guid></item></channel></rss>