<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: ModelForge</title><link>https://news.ycombinator.com/user?id=ModelForge</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 29 Jul 2026 05:05:03 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=ModelForge" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by ModelForge in "Kimi K3 Architecture Overview and Notes"]]></title><description><![CDATA[
<p>And adding to that, there is also the recurrent state in the Kimi Delta Attention. I wouldn't call it position information but more sth like "position sensitivity"</p>
]]></description><pubDate>Tue, 28 Jul 2026 22:56:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49091116</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=49091116</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49091116</guid></item><item><title><![CDATA[Kimi K3 Architecture Overview and Notes]]></title><description><![CDATA[
<p>Article URL: <a href="https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html">https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49085698">https://news.ycombinator.com/item?id=49085698</a></p>
<p>Points: 352</p>
<p># Comments: 56</p>
]]></description><pubDate>Tue, 28 Jul 2026 15:48:34 +0000</pubDate><link>https://sebastianraschka.com/blog/2026/kimi-k3-architecture-notes.html</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=49085698</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49085698</guid></item><item><title><![CDATA[Inkling: A New Open-Weight 975B Moe with a Few Surprises]]></title><description><![CDATA[
<p>Article URL: <a href="https://sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html">https://sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48934761">https://news.ycombinator.com/item?id=48934761</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Thu, 16 Jul 2026 14:03:05 +0000</pubDate><link>https://sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=48934761</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48934761</guid></item><item><title><![CDATA[Claude Code's Real Secret Sauce Isn't the Model]]></title><description><![CDATA[
<p>Article URL: <a href="https://sebastianraschka.com/blog/2026/claude-code-secret-sauce.html">https://sebastianraschka.com/blog/2026/claude-code-secret-sauce.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47588089">https://news.ycombinator.com/item?id=47588089</a></p>
<p>Points: 6</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 31 Mar 2026 14:43:51 +0000</pubDate><link>https://sebastianraschka.com/blog/2026/claude-code-secret-sauce.html</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=47588089</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47588089</guid></item><item><title><![CDATA[The State of LLMs 2025: Progress, Problems, and Predictions]]></title><description><![CDATA[
<p>Article URL: <a href="https://magazine.sebastianraschka.com/p/state-of-llms-2025">https://magazine.sebastianraschka.com/p/state-of-llms-2025</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46445076">https://news.ycombinator.com/item?id=46445076</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 31 Dec 2025 15:40:41 +0000</pubDate><link>https://magazine.sebastianraschka.com/p/state-of-llms-2025</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=46445076</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46445076</guid></item><item><title><![CDATA[A Researcher's Field Guide to Non-Standard LLM Architectures]]></title><description><![CDATA[
<p>Article URL: <a href="https://magazine.sebastianraschka.com/p/beyond-standard-llms">https://magazine.sebastianraschka.com/p/beyond-standard-llms</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45811756">https://news.ycombinator.com/item?id=45811756</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 04 Nov 2025 15:00:57 +0000</pubDate><link>https://magazine.sebastianraschka.com/p/beyond-standard-llms</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45811756</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45811756</guid></item><item><title><![CDATA[Explanation of Gated DeltaNet (Qwen3-Next and Kimi Linear)]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/08_deltanet/README.md">https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/08_deltanet/README.md</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45801296">https://news.ycombinator.com/item?id=45801296</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 03 Nov 2025 16:59:49 +0000</pubDate><link>https://github.com/rasbt/LLMs-from-scratch/blob/main/ch04/08_deltanet/README.md</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45801296</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45801296</guid></item><item><title><![CDATA[The Core Components of Modern LLMs and the Models Beyond Transformers [video]]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.youtube.com/watch?v=lONyteDR4XE">https://www.youtube.com/watch?v=lONyteDR4XE</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45722231">https://news.ycombinator.com/item?id=45722231</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 27 Oct 2025 15:40:25 +0000</pubDate><link>https://www.youtube.com/watch?v=lONyteDR4XE</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45722231</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45722231</guid></item><item><title><![CDATA[Popular Attention Alternatives: GQA, MLA, SWA]]></title><description><![CDATA[
<p>Article URL: <a href="https://sebastianraschka.com/llms-from-scratch/ch04/">https://sebastianraschka.com/llms-from-scratch/ch04/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45592961">https://news.ycombinator.com/item?id=45592961</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 15 Oct 2025 14:17:52 +0000</pubDate><link>https://sebastianraschka.com/llms-from-scratch/ch04/</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45592961</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45592961</guid></item><item><title><![CDATA[Multi-Head Latent Attention]]></title><description><![CDATA[
<p>Article URL: <a href="https://sebastianraschka.com/llms-from-scratch/ch04/05_mla/">https://sebastianraschka.com/llms-from-scratch/ch04/05_mla/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45571699">https://news.ycombinator.com/item?id=45571699</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 13 Oct 2025 18:24:28 +0000</pubDate><link>https://sebastianraschka.com/llms-from-scratch/ch04/05_mla/</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45571699</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45571699</guid></item><item><title><![CDATA[Thinking Machines Lab Co-Founder Departs for Meta]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.wsj.com/tech/ai/thinking-machines-lab-co-founder-departs-for-meta-442d7461">https://www.wsj.com/tech/ai/thinking-machines-lab-co-founder-departs-for-meta-442d7461</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45552181">https://news.ycombinator.com/item?id=45552181</a></p>
<p>Points: 7</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 11 Oct 2025 19:57:45 +0000</pubDate><link>https://www.wsj.com/tech/ai/thinking-machines-lab-co-founder-departs-for-meta-442d7461</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45552181</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45552181</guid></item><item><title><![CDATA[OpenAI's internal Slack messages could cost it billions in copyright suit]]></title><description><![CDATA[
<p>Article URL: <a href="https://sherwood.news/power/openais-internal-slack-messages-could-cost-them-billions-in-copyright-suit/">https://sherwood.news/power/openais-internal-slack-messages-could-cost-them-billions-in-copyright-suit/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45543544">https://news.ycombinator.com/item?id=45543544</a></p>
<p>Points: 8</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 10 Oct 2025 20:41:08 +0000</pubDate><link>https://sherwood.news/power/openais-internal-slack-messages-could-cost-them-billions-in-copyright-suit/</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45543544</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45543544</guid></item><item><title><![CDATA[LLM Evaluation from Scratch: Multiple Choice, Verifiers, Leaderboards, LLM Judge]]></title><description><![CDATA[
<p>Article URL: <a href="https://magazine.sebastianraschka.com/p/llm-evaluation-4-approaches">https://magazine.sebastianraschka.com/p/llm-evaluation-4-approaches</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=45482523">https://news.ycombinator.com/item?id=45482523</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 05 Oct 2025 15:55:26 +0000</pubDate><link>https://magazine.sebastianraschka.com/p/llm-evaluation-4-approaches</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=45482523</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45482523</guid></item><item><title><![CDATA[New comment by ModelForge in "Gemma 3 270M re-implemented in pure PyTorch for local tinkering"]]></title><description><![CDATA[
<p>No the compiled version is actually faster.<p>From that table, the A100 tok/sec (larger is faster) numbers are:<p>- Eager: 28<p>- Compiled: 128<p>And<p>- KV cache eager: 26<p>- KV cache compiled: 99<p>The reason that the KV cache is slower is likely because it's not GPU-optimized code. On CPU the KV cache is faster. To make it faster on GPU, you would pre-allocate the tensors on the device for example instead of `torch.cat`ting them on the fly</p>
]]></description><pubDate>Wed, 20 Aug 2025 20:48:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=44966290</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44966290</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44966290</guid></item><item><title><![CDATA[New comment by ModelForge in "Gemma 3 270M re-implemented in pure PyTorch for local tinkering"]]></title><description><![CDATA[
<p>Could be an artifact of the small size not fully taking advantage of the GPU. For example, for the slightly larger Qwen3 0.6B model the A100 is faster (you can see it when scrolling to the bottom here: <a href="https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/11_qwen3" rel="nofollow">https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/11...</a>)</p>
]]></description><pubDate>Wed, 20 Aug 2025 20:44:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=44966243</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44966243</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44966243</guid></item><item><title><![CDATA[New comment by ModelForge in "Gemma 3 270M re-implemented in pure PyTorch for local tinkering"]]></title><description><![CDATA[
<p>I'd say the common ones (besides educational) are<p>- private, on-device models (possibly with lower latency than models via web API); also edge devices<p>- algorithm research (faster and cheaper to prototype new ideas)<p>- cheap tasks, like classification/categorization; sure, you don't need a decoder-style LLM for that, but it has the advantage of being more free-form, which is useful in many scenarios; or maybe a sanity checker for grammar; or even a router to other model (GPT-5 style)</p>
]]></description><pubDate>Wed, 20 Aug 2025 20:39:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=44966190</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44966190</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44966190</guid></item><item><title><![CDATA[Gemma 3 270M re-implemented in pure PyTorch for local tinkering]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/12_gemma3">https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/12_gemma3</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44962059">https://news.ycombinator.com/item?id=44962059</a></p>
<p>Points: 417</p>
<p># Comments: 57</p>
]]></description><pubDate>Wed, 20 Aug 2025 14:01:26 +0000</pubDate><link>https://github.com/rasbt/LLMs-from-scratch/tree/main/ch05/12_gemma3</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44962059</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44962059</guid></item><item><title><![CDATA[New comment by ModelForge in "GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2"]]></title><description><![CDATA[
<p>I think GPT-4.5 was potentially the original GPT-5 model that was larger and pre-trained on more data. Too bad it was too expensive to deploy at scale so that we never saw the RL-ed version</p>
]]></description><pubDate>Sun, 10 Aug 2025 21:52:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=44858625</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44858625</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44858625</guid></item><item><title><![CDATA[New comment by ModelForge in "GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2"]]></title><description><![CDATA[
<p>The ollama one uses even less (around 13 GB), which is nice. Apparently the gpt-oss team also shared the mxfp4 optimizations for metal</p>
]]></description><pubDate>Sun, 10 Aug 2025 21:50:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=44858617</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44858617</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44858617</guid></item><item><title><![CDATA[New comment by ModelForge in "GPT-OSS vs. Qwen3 and a detailed look how things evolved since GPT-2"]]></title><description><![CDATA[
<p>Good point. LLMs lower the barrier to entry if someone has enough resources because those architectures are more robust to tweaks given one throws enough compute and data at them. You can even violate scaling laws and still get a good model (like Llama 3 showed back then)</p>
]]></description><pubDate>Sun, 10 Aug 2025 19:04:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=44857409</link><dc:creator>ModelForge</dc:creator><comments>https://news.ycombinator.com/item?id=44857409</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44857409</guid></item></channel></rss>