<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mailonce</title><link>https://news.ycombinator.com/user?id=mailonce</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 11 Sep 2026 16:56:48 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mailonce" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mailonce in "What happens when a GPU writes memory"]]></title><description><![CDATA[
<p>I've been running local image models on an older laptop recently, and memory behavior surprised me more than raw inference time.
One experiment briefly pushed private memory past 18 GB before I changed the allocation behavior. After terminating the worker/process, it dropped dramatically.
It made me realize how different "the model fits in memory" is from "the whole inference pipeline behaves well in memory."</p>
]]></description><pubDate>Thu, 10 Sep 2026 21:32:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49650436</link><dc:creator>mailonce</dc:creator><comments>https://news.ycombinator.com/item?id=49650436</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49650436</guid></item></channel></rss>