<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: kbwal7</title><link>https://news.ycombinator.com/user?id=kbwal7</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 06 Sep 2026 07:41:39 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=kbwal7" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by kbwal7 in "One geometric idea that explains SGD, AdamW, and Muon"]]></title><description><![CDATA[
<p>I'm honestly not too aware of the true scene of which optimizers are used nowadays by which labs (though I think Moonshot has shown Muon's ability to scale to large models, and Muon is definitely more used now). 
That being said, obviously orthogonalization is pretty expensive! I looked it up and it's roughly 5-15% slower per step, but it converges in fewer steps fwiw. I'm guessing it's also less battle-tested than AdamW is, so AdamW still might be the safer option (and you're gonna have to use AdamW for your 1D params anyways).</p>
]]></description><pubDate>Fri, 04 Sep 2026 19:44:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49569320</link><dc:creator>kbwal7</dc:creator><comments>https://news.ycombinator.com/item?id=49569320</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49569320</guid></item><item><title><![CDATA[One geometric idea that explains SGD, AdamW, and Muon]]></title><description><![CDATA[
<p>Article URL: <a href="https://kbwal.github.io/writing/notes-on-muon/">https://kbwal.github.io/writing/notes-on-muon/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49569030">https://news.ycombinator.com/item?id=49569030</a></p>
<p>Points: 6</p>
<p># Comments: 4</p>
]]></description><pubDate>Fri, 04 Sep 2026 19:21:49 +0000</pubDate><link>https://kbwal.github.io/writing/notes-on-muon/</link><dc:creator>kbwal7</dc:creator><comments>https://news.ycombinator.com/item?id=49569030</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49569030</guid></item><item><title><![CDATA[New comment by kbwal7 in "Corporate America is getting hooked on open-source AI"]]></title><description><![CDATA[
<p>I do think there are now open weight models that are on par with (or beating) Opus 4.5 by now (e.g. Kimi K3, GLM5.3). But yeah obviously the frontier closed source models seem to have pulled away once again, so open weight seems to be a few months behind right now (which might be too long to wait for a lot of people!).</p>
]]></description><pubDate>Fri, 04 Sep 2026 16:04:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49566528</link><dc:creator>kbwal7</dc:creator><comments>https://news.ycombinator.com/item?id=49566528</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49566528</guid></item><item><title><![CDATA[New comment by kbwal7 in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Olmo is one truly open source model. <a href="https://allenai.org/blog/olmo3" rel="nofollow">https://allenai.org/blog/olmo3</a></p>
]]></description><pubDate>Fri, 04 Sep 2026 15:53:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49566387</link><dc:creator>kbwal7</dc:creator><comments>https://news.ycombinator.com/item?id=49566387</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49566387</guid></item><item><title><![CDATA[New comment by kbwal7 in "IBM Bob"]]></title><description><![CDATA[
<p>I think it's just a harness that uses proprietary (or maybe IBM is hosting open weight?) models. They say on the site (in the FAQs) that you can't even choose which model!<p>So it seems to be a harness that auto-routes your queries with no way to change this.</p>
]]></description><pubDate>Fri, 04 Sep 2026 15:43:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49566253</link><dc:creator>kbwal7</dc:creator><comments>https://news.ycombinator.com/item?id=49566253</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49566253</guid></item><item><title><![CDATA[New comment by kbwal7 in "Marin 32B Spike Fixing"]]></title><description><![CDATA[
<p>It's pretty incredible how robust neural networks are to even things like architectural changes mid-run. The team at Meta when training their own copy of GPT-3 (called OPT-3) even changed optimizers mid-run (from AdamW -> SGD -> AdamW, <a href="https://arxiv.org/pdf/2205.01068" rel="nofollow">https://arxiv.org/pdf/2205.01068</a> section 2.5)!</p>
]]></description><pubDate>Fri, 04 Sep 2026 08:06:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49561881</link><dc:creator>kbwal7</dc:creator><comments>https://news.ycombinator.com/item?id=49561881</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49561881</guid></item></channel></rss>