<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: versteegen</title><link>https://news.ycombinator.com/user?id=versteegen</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 09 Oct 2026 07:48:04 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=versteegen" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by versteegen in "Claude Haiku 5.5"]]></title><description><![CDATA[
<p>You mean if you used a modern LLM of GPT-2-level quality. Vanilla transformers like GPT 2 are ridiculously inefficient in comparison.</p>
]]></description><pubDate>Wed, 07 Oct 2026 23:23:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=50000121</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=50000121</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50000121</guid></item><item><title><![CDATA[New comment by versteegen in "When did Google get so weird?"]]></title><description><![CDATA[
<p>You're asking for something unreasonable. The number of Google searches per day is enormous and they haven't even been able to roll out AI overviews to everyone yet (they're missing in a new Firefox profile I just created). I wouldn't be surprised if the free tier frontier models cost over 100x more to serve than the AI overviews.</p>
]]></description><pubDate>Sun, 27 Sep 2026 23:10:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49871748</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49871748</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49871748</guid></item><item><title><![CDATA[New comment by versteegen in "When did Google get so weird?"]]></title><description><![CDATA[
<p>It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)<p><pre><code>  TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
  next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
  pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
</code></pre>
<a href="https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-do-a-concerning-amount-with-no-chain-of-thought" rel="nofollow">https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...</a></p>
]]></description><pubDate>Sun, 27 Sep 2026 23:02:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49871690</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49871690</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49871690</guid></item><item><title><![CDATA[New comment by versteegen in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>Ugh, people are still saying the Codex limits are more generous. They're not, Claude's are over 2x higher, have been for months! [1] It's just that Claude uses far more tokens, 2-3x is common. Except sometimes GPT will use just as many or even go into a compact loop and then your quota is gone, little headroom for hard tasks.<p>[1] <a href="https://devforth.io/agents-for-code/?sortby=monthly-value" rel="nofollow">https://devforth.io/agents-for-code/?sortby=monthly-value</a> And I can confirm the numbers, I subscribe to both and watch the numbers</p>
]]></description><pubDate>Tue, 01 Sep 2026 18:49:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49526284</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49526284</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49526284</guid></item><item><title><![CDATA[New comment by versteegen in "Nearly 3M Teslas recalled in China over hidden door handles"]]></title><description><![CDATA[
<p>> Ok the rear red turn signal lights were already a trend in the US before Tesla. But please reverse this too.<p>I'm surprised. But it shouldn't be up to Tesla; it's illegal and they never get on the road in countries with proper regulations.</p>
]]></description><pubDate>Mon, 24 Aug 2026 05:30:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49415503</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49415503</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49415503</guid></item><item><title><![CDATA[New comment by versteegen in "Compression is prediction"]]></title><description><![CDATA[
<p>I urge you to reconsider your beliefs. You are missing something important because you are thinking in terms of low-dimensional statistics. Deep learning doesn't just fit data, it finds features (abstractions) of the data.</p>
]]></description><pubDate>Sat, 15 Aug 2026 00:36:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49306334</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49306334</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49306334</guid></item><item><title><![CDATA[New comment by versteegen in "Qwen 3.8 27B"]]></title><description><![CDATA[
<p>IME using 5.6 Luna and DS V4 Flash, I notice that although they are excellent at programming, even Opus-like in the way they try to debug, the thing they are worst at is inferring user intent and making good decisions with little information. They are absolutely <i>terrible</i> at that, will misinterpret small wording ambiguities. I suspect that's an ability you can't add with RL training, that it requires the depth of understanding from vast pre-training.</p>
]]></description><pubDate>Fri, 14 Aug 2026 15:46:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49300365</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49300365</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49300365</guid></item><item><title><![CDATA[New comment by versteegen in "DeepSeek V4 Pro 0813"]]></title><description><![CDATA[
<p>I think this is the best and most useful way to measure model intelligence. In my experience it's what really sets apart the capable models from the best. A small model can be RL trained to be extremely good at programming or narrow problem solving for its size (eg 5.6 Luna, DS4 Flash, Qwen 3.6 27B), but even Luna is IME comparatively awful at understanding intent and making good decisions with limited guidance.</p>
]]></description><pubDate>Thu, 13 Aug 2026 03:22:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49281430</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49281430</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49281430</guid></item><item><title><![CDATA[New comment by versteegen in "Compression is prediction"]]></title><description><![CDATA[
<p>If doesn't correspond cleanly. I can see why you draw the link, because LZ compression will replace words with symbols but BPE is a non-contextual entropy encoding while LZ is contextual and adaptive and that makes it very different. I think BPE actually has more in common with Huffman encoding.</p>
]]></description><pubDate>Wed, 12 Aug 2026 01:52:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49266970</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49266970</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49266970</guid></item><item><title><![CDATA[New comment by versteegen in "Compression is prediction"]]></title><description><![CDATA[
<p>That conception of knowledge is interesting, but I think using the label 'knowledge' for it is very problematic, it's too far from common definitions. The fact that you have to carve out an exception for mathematics already shows there's a problem. Because if maths, shouldn't thought experiments also produce new knowledge? You're excluding special and general relativity. It seems to me that what the concept actually describes is "information about the world".</p>
]]></description><pubDate>Wed, 12 Aug 2026 01:50:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49266948</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49266948</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49266948</guid></item><item><title><![CDATA[New comment by versteegen in "Compression is prediction"]]></title><description><![CDATA[
<p>There is a distinction between a compressor for a fixed dataset and one for an unknown population from which we have a sample. The optimal compressor for the sample may be the single best guess for the population, but that's not what Solomonoff induction does. It begins with a prior that allows all possible programs, and it never assigns all probability to the single optimal compressor, so it has no problem with the all-zeroes example.<p>But the Hutter prize (of which I'm a big fan) is for ever-more-optimal compressors, and in fact many of the solutions don't generalise to other input data without stripping out various tricks.</p>
]]></description><pubDate>Wed, 12 Aug 2026 01:33:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49266834</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49266834</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49266834</guid></item><item><title><![CDATA[New comment by versteegen in "Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs"]]></title><description><![CDATA[
<p>> For me it’s the massive amount of resources it takes to produce and run one<p>It's amazing that LLM pretraining is both extremely data inefficient at learning concepts and cognitive functions from the training data compared to humans, while actually being quite efficient at learning facts, memorising things seen just a few times.<p>I used to likewise think that the resources required to run large transformers were absurd, but the architectures are far more efficient now than 3 years ago and I underestimated just massive the parallelisation advantage of transformers is, how many TFLOPS effective you can get. You can already run amazingly decent LLMs on PCs and phones.<p>I generally agree with you, but my view has shifted from "we need to augment or replace LLMs" to it there being far more efficient algorithms possible but it not actually being necessary for fulfilling most goals.</p>
]]></description><pubDate>Fri, 07 Aug 2026 02:27:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49205272</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49205272</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49205272</guid></item><item><title><![CDATA[New comment by versteegen in "Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs"]]></title><description><![CDATA[
<p>You misread. "Pain and suffering" not "death". Of all the pain and suffering in the world, a vast amount of it really is our own fault. Famines and wars shouldn't happen. And if you see a country border with poverty and one side and prosperity on the other you can't say that was the only possibility.</p>
]]></description><pubDate>Fri, 07 Aug 2026 01:57:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49205099</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49205099</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49205099</guid></item><item><title><![CDATA[New comment by versteegen in "Connes' Rigidity Theorem: Disproof of Open AI's Counterexample and Proof"]]></title><description><![CDATA[
<p>To save anyone else the trouble: discussion there is not really worth looking at (largely a flame war), except: the author of this disproof seems to be a crank, and the disproof's been refuted.</p>
]]></description><pubDate>Mon, 03 Aug 2026 02:42:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49150642</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49150642</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49150642</guid></item><item><title><![CDATA[New comment by versteegen in "Kimi K3-256k"]]></title><description><![CDATA[
<p>...but we're talking about compaction, and opencode's compaction is (or was) terrible. I've seen so many horrible problems that I keep it disabled (with an envvar flag, because even the config flag to turn it off was broken).</p>
]]></description><pubDate>Thu, 30 Jul 2026 07:56:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49107159</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49107159</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49107159</guid></item><item><title><![CDATA[New comment by versteegen in "Pacing the frontier"]]></title><description><![CDATA[
<p>Wow, remarkable. Clearly Aum is a different league from the lone-wolf "Fort Detrick guy", treat the risks separately. But I'll take these questions as rhetorical. (See my reply to the sibling comment.) I can't answer them and I'm not defending the guardrails on Claude; in their current form I too find them pretty ridiculous.</p>
]]></description><pubDate>Wed, 29 Jul 2026 07:43:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49094508</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49094508</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49094508</guid></item><item><title><![CDATA[New comment by versteegen in "Pacing the frontier"]]></title><description><![CDATA[
<p>I didn't argue "must be regulated". I'm arguing AI is powerful (at achieving things, and hence has dual-use dangers). Many people won't even admit that, which is the part that really annoys me: they don't even want to have a conversation about risks because somehow AI is not actually a powerful general-purpose tool. Of course everything you said is true, though many of those things aren't very powerful. But the internet and social media and smartphones certainly have been of great utility for terrorism and child exploitation (and probably also causing many would-be-terrorists to get picked up).</p>
]]></description><pubDate>Wed, 29 Jul 2026 06:47:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49094152</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49094152</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49094152</guid></item><item><title><![CDATA[New comment by versteegen in "Pacing the frontier"]]></title><description><![CDATA[
<p>Thank you for taking this seriously enough to write this, and anyone else likewise.<p>But this is attacking a strawman, amateur bioterrorists. AI is a force multiplier in the hands of an expert. If it took a team before, maybe a single malicious actor can accomplish it now that AI can fill in the parts they aren't well-versed in. And that dramatically increases the chance of it happening.<p>Being at risk of killing yourself also just makes success X% less likely, but if X < 90 that doesn't mitigate much.</p>
]]></description><pubDate>Wed, 29 Jul 2026 02:17:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49092689</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49092689</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49092689</guid></item><item><title><![CDATA[New comment by versteegen in "Pacing the frontier"]]></title><description><![CDATA[
<p>The old observation that people tend to define AI as whatever computers can't do yet is as true as ever. It's getting a bit absurd, moving from demanding "general intelligence" to replicating human cognitive phenomenology (the experience of cognitive activities). Yet LLMs can already somewhat (confabulation-prone) introspect their own "internal unspoken thoughts" in their residual streams, quite fascinating.</p>
]]></description><pubDate>Wed, 29 Jul 2026 01:44:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49092452</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49092452</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49092452</guid></item><item><title><![CDATA[New comment by versteegen in "ARC-AGI Leaderboard"]]></title><description><![CDATA[
<p>Isn't it ~$3000 per week? Extrapolating from the current limit on Pro plans.</p>
]]></description><pubDate>Sat, 25 Jul 2026 11:43:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49046763</link><dc:creator>versteegen</dc:creator><comments>https://news.ycombinator.com/item?id=49046763</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49046763</guid></item></channel></rss>