<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mochizou</title><link>https://news.ycombinator.com/user?id=mochizou</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 07 Oct 2026 01:50:12 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mochizou" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mochizou in "Dust: Pretraining Transformers Without Backpropagation"]]></title><description><![CDATA[
<p>Maybe the practical path is to separate fast-changing memory from slow-changing weights. Most things an agent learns during use probably don't need to become parameters immediately.</p>
]]></description><pubDate>Tue, 06 Oct 2026 11:29:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49976954</link><dc:creator>mochizou</dc:creator><comments>https://news.ycombinator.com/item?id=49976954</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49976954</guid></item><item><title><![CDATA[New comment by mochizou in "I built non-autoregressive decision models with RL a year ago"]]></title><description><![CDATA[
<p>I tried Jev a bit, and the zero-shot performance already felt good enough to be useful. If you find the right place for that tradeoff, I don’t really see why “you could fine-tune BERT” is much of a criticism.</p>
]]></description><pubDate>Sun, 20 Sep 2026 14:37:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49776324</link><dc:creator>mochizou</dc:creator><comments>https://news.ycombinator.com/item?id=49776324</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49776324</guid></item></channel></rss>