<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: dnhkng</title><link>https://news.ycombinator.com/user?id=dnhkng</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 20 Sep 2026 14:07:25 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=dnhkng" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by dnhkng in "DeepSeek-V4-Flash Update"]]></title><description><![CDATA[
<p>Totally! This with DwarfStar delivers <i>usable</i> local AI (I hope!)</p>
]]></description><pubDate>Fri, 31 Jul 2026 06:41:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49119770</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=49119770</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49119770</guid></item><item><title><![CDATA[New comment by dnhkng in "DeepSeek-V4-Flash Update"]]></title><description><![CDATA[
<p>I think the commenter means the Flash vs Terra benchmarks.</p>
]]></description><pubDate>Fri, 31 Jul 2026 06:40:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49119764</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=49119764</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49119764</guid></item><item><title><![CDATA[New comment by dnhkng in "DeepSeek-V4-Flash Update"]]></title><description><![CDATA[
<p>It will be fair if they release the harness though. I think now the future will be paired model-harness releases, not just weight dumps.<p>The performance changes are so big with the right harness that is makes sense to engineer the harness and fine-tune the model to one another from the start.</p>
]]></description><pubDate>Fri, 31 Jul 2026 06:39:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49119760</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=49119760</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49119760</guid></item><item><title><![CDATA[New comment by dnhkng in "DeepSeek-V4-Flash Update"]]></title><description><![CDATA[
<p>DeepSeek V4 Flash (Preview → 2026-07-31)<p>• Terminal Bench: 56.9 → 82.7 (+25.8)<p>• Toolathlon: 51.8 → 70.3 (+18.5)<p>Compared to GPT-5.6 Terra:<p>• Terminal Bench: Flash 82.7 vs Terra 78.4<p>• Toolathlon: Flash 70.3 vs Terra 53.1<p>• DeepSWE: Flash 54.4 vs Terra 69.6<p>• Agents' Last Exam: Flash 25.2 vs Terra 50.4<p>Trading blows with Terra, which is pretty interesting. No clear winner on these benchmarks, and wildy differeing scores. Very interesting!</p>
]]></description><pubDate>Fri, 31 Jul 2026 06:16:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49119608</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=49119608</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49119608</guid></item><item><title><![CDATA[DeepSeek-V4-Flash Update]]></title><description><![CDATA[
<p>Article URL: <a href="https://api-docs.deepseek.com/updates/">https://api-docs.deepseek.com/updates/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49119559">https://news.ycombinator.com/item?id=49119559</a></p>
<p>Points: 745</p>
<p># Comments: 347</p>
]]></description><pubDate>Fri, 31 Jul 2026 06:08:36 +0000</pubDate><link>https://api-docs.deepseek.com/updates/</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=49119559</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49119559</guid></item><item><title><![CDATA[New comment by dnhkng in "Claude Fable produced a counterexample to the Jacobian Conjecture"]]></title><description><![CDATA[
<p>Maybe instead of proofs, we should encourage the publishing of prompts and reasoning traces?</p>
]]></description><pubDate>Mon, 20 Jul 2026 07:26:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48975355</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=48975355</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48975355</guid></item><item><title><![CDATA[New comment by dnhkng in "Claude Fable produced a counterexample to the Jacobian Conjecture"]]></title><description><![CDATA[
<p>Maybe instead of proofs, we should encourage the publishing of prompts and reasoning traces?</p>
]]></description><pubDate>Mon, 20 Jul 2026 07:26:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=48975354</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=48975354</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48975354</guid></item><item><title><![CDATA[New comment by dnhkng in "A global workspace in language models"]]></title><description><![CDATA[
<p>Author here:<p>Yeah, the encoder and decoder stuff is explicit, but the internal structure in generated during training. I don't think the big labs were doing this back when I did the research; no one was back in '24.<p>I just didn't get round to publishing for years, because I have a day job.<p>By the way, it still works! I tested it earlier this year on Qen3.6 and you still see improvements, so either a) no one actually paid attention, or b) it has more room to scale.</p>
]]></description><pubDate>Tue, 07 Jul 2026 14:15:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48818215</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=48818215</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48818215</guid></item><item><title><![CDATA[New comment by dnhkng in "A global workspace in language models"]]></title><description><![CDATA[
<p>Too bad they also don't give anything back to individual researchers. Oh well, wasn't expecting much.</p>
]]></description><pubDate>Tue, 07 Jul 2026 13:02:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48817177</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=48817177</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48817177</guid></item><item><title><![CDATA[New comment by dnhkng in "Show HN: I built a hardware quantum RNG and wired it into a Magic 8-Ball"]]></title><description><![CDATA[
<p>I have added an endpoint for CLI-dwellers:<p>curl -X POST <a href="https://quantumlever.stream/api/magic-8-ball" rel="nofollow">https://quantumlever.stream/api/magic-8-ball</a> \
  -H "Content-Type: application/json" \
  -d '{"question":"Will my sampler touch the multiverse?"}'</p>
]]></description><pubDate>Sat, 27 Jun 2026 12:32:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=48697686</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=48697686</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48697686</guid></item><item><title><![CDATA[Show HN: I built a hardware quantum RNG and wired it into a Magic 8-Ball]]></title><description><![CDATA[
<p>Gday, author here!<p>I've wanted to hack together a "real" quantum random number generator for another upcoming project, and I got carried away a bit, and went down the 'over-engineering' cliff.  So, for your nerdy enjoyment, I have documented it all up, and I added something cool for fellow "Multiple World Interpretation" followers in the Quantum Mechanics debate.<p>This QRNG uses sexy bits:  Each is the decision of a photon to go left or right after hitting a 50:50 beam splitter.  Standard kinda device, where you attenuate a light source down to single photons, offer them semi-mirror to bounce off, and see which PMT detector they hit (or which universe we ended up in ;) ). Basically, Through → bit=0. Bounce → bit=1.<p>As I take the MWI interpretation of Quantum Mechanics (its the more fun options), I have also built a Quantum Magic 8-Ball. Ask it a questions, and you get and receive exactly one answer here, plus every possible answer across the multiverse.<p><a href="https://quantumlever.stream/oracle" rel="nofollow">https://quantumlever.stream/oracle</a><p>Enjoy!</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48689891">https://news.ycombinator.com/item?id=48689891</a></p>
<p>Points: 13</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 26 Jun 2026 18:05:57 +0000</pubDate><link>https://dnhkng.github.io/posts/building-the-beam-universe-splitter/</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=48689891</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48689891</guid></item><item><title><![CDATA[New comment by dnhkng in "Do LLMs Break the Sapir-Whorf Hypothesis?"]]></title><description><![CDATA[
<p>Yes, dammit.<p>Author here.<p>I drafted it before I left for holiday, at it's not ready to publish.<p>It wasn't supposed to be officially posted yet, but I ran out of time before my flight.<p>My apologies!</p>
]]></description><pubDate>Thu, 02 Apr 2026 04:24:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=47609971</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47609971</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47609971</guid></item><item><title><![CDATA[New comment by dnhkng in "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?"]]></title><description><![CDATA[
<p>Thanks!<p>I have pushed basic code to GitHub (<a href="https://github.com/dnhkng/RYS" rel="nofollow">https://github.com/dnhkng/RYS</a>)<p>Some interesting areas to explore might be a combination of deleting some layers and duplicating others. i.e. reduce VRAM by dropping some layer (this works, well documented), and recovering performance by duplicating others (saves VRAM). I am not pursuing this, but it seems interesting!</p>
]]></description><pubDate>Tue, 24 Mar 2026 17:21:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=47506104</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47506104</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47506104</guid></item><item><title><![CDATA[New comment by dnhkng in "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?"]]></title><description><![CDATA[
<p>Author here: The code is up on GitHub.<p>The probes I used seem to help identify good configurations, but are quite noisey. A small probe set was initially used to make the scan tractable, and then the higher ranked models were retested on a set ~10x larger.</p>
]]></description><pubDate>Tue, 24 Mar 2026 14:56:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=47503576</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47503576</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47503576</guid></item><item><title><![CDATA[New comment by dnhkng in "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?"]]></title><description><![CDATA[
<p>Author here: That was done in this blog post, in the beam search. I started with the best re-layer configs, and iteratively added more blocks, including the same multiple times, during a long beam search.<p>It turns out this does not help (somewhat surprisingly).</p>
]]></description><pubDate>Tue, 24 Mar 2026 14:54:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=47503539</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47503539</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47503539</guid></item><item><title><![CDATA[New comment by dnhkng in "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?"]]></title><description><![CDATA[
<p>There was some work done on this a while back, during the FrankenMerge craze of 23'<p>I am working with TurboDerp to integrate this into the Exllama v3 format.</p>
]]></description><pubDate>Tue, 24 Mar 2026 13:59:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=47502721</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47502721</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47502721</guid></item><item><title><![CDATA[New comment by dnhkng in "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?"]]></title><description><![CDATA[
<p>Author here. Another thing I want to highlight: the language-agnostic "thinking space" finding came from Evan Maunder, who read Part 1 and ran an elegant experiment — same sentence in English, Mandarin, and Base64, cosine similarity at every layer. The representations converge by the early layers, stay nearly identical through the mid-stack, then diverge again at the end as the model commits to an output format.<p>I extended this to a 2×2 design (two languages × two content types) and the result is even starker: by layer 10, cross-language same-content pairs are more similar than same-language different-content pairs. The model cares about what you're saying, not what language you're saying it in.<p>This is also what makes layer duplication work — those mid-stack layers operate in a space where input and output distributions match, so you can loop through them without breaking anything. The encoding and decoding boundaries are where the blue walls show up in the heatmaps.</p>
]]></description><pubDate>Tue, 24 Mar 2026 13:56:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=47502672</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47502672</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47502672</guid></item><item><title><![CDATA[New comment by dnhkng in "LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language?"]]></title><description><![CDATA[
<p>Author here. The result that surprised me most: after evaluating 3,024 beam search candidates, training a surrogate model on ~4,600 measurements, and scoring 2 million configurations — the Pareto-optimal configs were all simple contiguous blocks. No exotic multi-block compositions, no sparse repeats. Just "repeat layers 31–33" and you're on the efficiency frontier.<p>I think this says something interesting about how transformers organise computation internally. The mid-stack reasoning circuits are coherent enough that you can loop through them twice without distribution mismatch. The encoding/decoding boundaries are not.</p>
]]></description><pubDate>Tue, 24 Mar 2026 13:30:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=47502311</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47502311</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47502311</guid></item><item><title><![CDATA[Show HN: More LLM Neuroanatomy: A Hint of a Universal Language?]]></title><description><![CDATA[
<p>Article URL: <a href="https://dnhkng.github.io/posts/rys-ii/">https://dnhkng.github.io/posts/rys-ii/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47502295">https://news.ycombinator.com/item?id=47502295</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 24 Mar 2026 13:29:04 +0000</pubDate><link>https://dnhkng.github.io/posts/rys-ii/</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47502295</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47502295</guid></item><item><title><![CDATA[New comment by dnhkng in "Show HN: How I topped the HuggingFace open LLM leaderboard on two gaming GPUs"]]></title><description><![CDATA[
<p>I stick with models I can run on VRAM, but DeepSeek Speciale have the best reasoning capabilities of the models I can actually run (<a href="https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale" rel="nofollow">https://huggingface.co/deepseek-ai/DeepSeek-V3.2-Speciale</a>).  What hardware can you access?<p>I have Deepseek etc, but inferencing on DDR5 would take about 2-3 weeks for a simple scan.  I think this works best with dense models, but it also seems ok with MoE.<p>@everyone: Can someone hook me up with Nvidia sponsorship?</p>
]]></description><pubDate>Fri, 13 Mar 2026 16:41:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=47366673</link><dc:creator>dnhkng</dc:creator><comments>https://news.ycombinator.com/item?id=47366673</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47366673</guid></item></channel></rss>