<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: SerdarGl</title><link>https://news.ycombinator.com/user?id=SerdarGl</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 08 Oct 2026 05:13:37 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=SerdarGl" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by SerdarGl in "Dust: Pretraining Transformers Without Backpropagation"]]></title><description><![CDATA[
<p>At massive scales 0th order methods will parallelize better than backprop especially along depth, u can train very deep models pipeline parallel without bubbles</p>
]]></description><pubDate>Tue, 06 Oct 2026 02:23:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49973430</link><dc:creator>SerdarGl</dc:creator><comments>https://news.ycombinator.com/item?id=49973430</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49973430</guid></item></channel></rss>