<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: gwerbin</title><link>https://news.ycombinator.com/user?id=gwerbin</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 19 Aug 2026 01:53:43 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=gwerbin" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by gwerbin in "The Benchmarkpocalypse"]]></title><description><![CDATA[
<p>If you watch the thinking traces of just about any modern LLM, you might be surprised at how much "uncertainty" is in there. Weak models with no thinking limits vacillate back-and-forth back-and-forth on a topic for potentially thousands of tokens before gradually spiraling towards some kind of an answer. Which makes it all the more interesting that "I don't know" is so rarely the final prediction, even with so much waffling in the chain of thought.<p>Until the big labs decide to start adding synthetic "I don't know" outcomes to their data sets, I've been thinking that the best way to evaluate uncertainty is to have a separate LLM monitoring the conversation and asking it to classify if the agent is overstating its confidence. On the other hand I've also noticed that most models will tell you they don't know something if you specifically include it in the prompt, eg "if you don't know the answer,  just say so" and/or "be clear about any gaps in your knowledge that would reduce the confidence of your response" etc. but even with the big frontier models I have noticed some quality degradation if I throw too many instructions into the system prompt. I have a little more faith in harness-level engineering than in praying to the token generation gods.<p>That said, there is a completely different form of "uncertainty" in which the LLM tends to place very high trust in its own prior outputs as well as user provided inputs. Again if you look at the thinking traces, these models will try very very hard to rationalize the inputs they are given, falling back to the possibility of user error only after working through several alternative possibilities, maybe even investigating data or source code in the process. And if your context is big enough, the model might just completely miss when pieces of information conflict.</p>
]]></description><pubDate>Tue, 18 Aug 2026 23:18:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49354248</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49354248</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49354248</guid></item><item><title><![CDATA[New comment by gwerbin in "The Benchmarkpocalypse"]]></title><description><![CDATA[
<p>The only people who think LLMs would make good lawyers are the people selling LLMs. The more practical among us recognize that LLMs are our amazing tools for searching through and making sense of large amounts of text with a high level of sophistication, which can significantly enhance the productivity of a human lawyer.</p>
]]></description><pubDate>Tue, 18 Aug 2026 23:09:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49354136</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49354136</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49354136</guid></item><item><title><![CDATA[New comment by gwerbin in "The Benchmarkpocalypse"]]></title><description><![CDATA[
<p>This is a "don't make me tap the sign" moment. LLMs are next token prediction models. If there are factual errors, confused ideas, etc. in the preceding tokens, that will affect the generation of subsequent tokens, and the error accumulates.<p>Case in point, I hit an error in a SQL query today because it turned out I was trying to do something that wasn't supported by the query engine. I pasted the error message and a bit of background info into my Claude Code session with Sonnet 5 High, it worked on a response for an unexpectedly long amount of time, including consulting the advisor model, and then came back with an explanation of the mistake I made in my query. Except it turned out I pointed it to the wrong file, and there wasn't a mistake in that file. It had completely taken for granted that the pasted error output was a real error and went on some wild goose chase.<p>Part of why the current gen models feel so smart is that they're getting better (via CoT and training) at recognizing when something is wrong and then back up to reassess. So it's easy to forget that it really is just token prediction, and (pending the next big advancement) there's only so much you can do with that.</p>
]]></description><pubDate>Tue, 18 Aug 2026 07:45:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49342677</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49342677</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49342677</guid></item><item><title><![CDATA[New comment by gwerbin in "Show HN: Eigendrum - Draw any shape and hear what it sounds like as a drum"]]></title><description><![CDATA[
<p>> Pitch here is a reference: the note a circle of this area would sound. Tension and density are yours, and so is fade, which is material and air. Where each shape's own fundamental lands above that reference is not yours, and neither is any overtone ratio. Both are computed.<p>You might want to do a human editing pass on the copy. This reads like it came right out of your Claude Code session.</p>
]]></description><pubDate>Sat, 15 Aug 2026 18:45:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49313172</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49313172</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49313172</guid></item><item><title><![CDATA[New comment by gwerbin in "The other Sean Byrne doesn't exist"]]></title><description><![CDATA[
<p>Government entities in the USA already work like this. That's how all this came about.</p>
]]></description><pubDate>Sat, 15 Aug 2026 18:40:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49313129</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49313129</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49313129</guid></item><item><title><![CDATA[New comment by gwerbin in "The other Sean Byrne doesn't exist"]]></title><description><![CDATA[
<p>Also because there's no real due process associated with these various lists, they can be used to punish political enemies.</p>
]]></description><pubDate>Sat, 15 Aug 2026 18:35:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49313073</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49313073</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49313073</guid></item><item><title><![CDATA[New comment by gwerbin in "When Genius Fails: The Intellectual Arrogance of the AI Labs"]]></title><description><![CDATA[
<p>I'm not saying we won't have breakthroughs or significant capability changes. I'm saying that any reasonable expected rate of breakthroughs is not enough to maintain exponential R&D effort <i>and</i> to have it translate into exponential progress towards AGI, unless it's an absolutely monumental discovery on par with the GPT LLM itself. It would be impossible to predict, and it would be totally fallacious to credit this author with being a visionary if such a discovery does in fact arise.</p>
]]></description><pubDate>Sat, 15 Aug 2026 04:00:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49307519</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49307519</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49307519</guid></item><item><title><![CDATA[New comment by gwerbin in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>N=2 anecdata but just this week we were discussing  setting up a couple of seats with OpenAI as a trial for switching. There are other advantages too, such as being able to bring your own harness including Ai-integrated editors / ACP clients such as Jetbrains, VS Code, and Zed. I think OpenAI and Altman are a clear step more evil than Anthropic and Amodei so I really hate to say it, but with the degradation in model output interpretability, all of the cleverness and power of the Claude Code harness hasn't been enough to offset a genuine falloff in productivity for anything other than total hands-off automation.<p>That said, the duo of Opus 5 and Sonnet 5 do a fantastic job at fully automated work, and Claude Code still stands head and shoulders above the rest.</p>
]]></description><pubDate>Sat, 15 Aug 2026 01:56:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49306874</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49306874</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49306874</guid></item><item><title><![CDATA[New comment by gwerbin in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>They want you to use Sonnet to explain what Opus is trying to say. They're not optimizing for token efficiency.</p>
]]></description><pubDate>Fri, 14 Aug 2026 16:14:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49300783</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49300783</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49300783</guid></item><item><title><![CDATA[New comment by gwerbin in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>I think it's a deliberate steganography choice. You can spot Claude vocabulary a mile away, which maybe means you can spot distillations a mile away.<p>But I agree, the GPT models are so much simpler to work with, they have so much less personality and fewer quirks. They also are a little less aggressive about triple checking every little assumption immediately in a stack of 30 tool calls (but I haven't used 5.6 Sol yet so maybe that's not true anymore).</p>
]]></description><pubDate>Fri, 14 Aug 2026 16:11:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49300741</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49300741</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49300741</guid></item><item><title><![CDATA[New comment by gwerbin in "When Genius Fails: The Intellectual Arrogance of the AI Labs"]]></title><description><![CDATA[
<p>Wait, is this the group who predicted that we'd have expontential AI progress and AGI in 2027 because we'd have AIs training AIs? Predicated on the preposterous assumption that R&DEffort == RateOfProgress. The more sober and realistic assessment is that R&DEffort >= RateOfProgress, with only short bursts of achieving the upper bound. Leave it to a brilliant 25 year old to overfit to one of those short bursts of progress and then assume it will continue forever just as it always has.<p>What actually happened is that AI-assisted training of LLMs (mind you -- stuck on that same old transformer architecture, with nothing resembling a medium-term memory, and fine-tuning still doesn't really work as a form of "learning") is right now giving us ~linear-ish progress... beecause we already hit the slowdown in the S curve of what human researchers can achieve.<p>We are also energy-constrained in ways that the exponential forecasts have no answer for other. Compute is not getting more efficient fast enough to achieve anything close to exponential growth driven by growth in computing power. There are social, moral, and political limits on the amount of energy we can dedicate to AI training in any given time period. And it's not just energy, we're short on RAM wafers and water for cooling and eventually we're going to hit other limits on various minerals and elements of the supply chain etc. etc. So <i>even if</i> it were true that R&DEffort == RateOfProgress without resource constraints, we're likely not going to see that exponential progress because we also need corresponding exponential scifi-scale progress in computing efficiency, cooling, and, energy delivery.<p>Basically this kid made a bunch of very clever and grand predictions which were always kind of ridiculous and no, they have not come to pass in 2026 and we are not at all on track to achieve AGI in the next 12 months, despite what Altman keeps trying to make the public think.</p>
]]></description><pubDate>Fri, 14 Aug 2026 15:48:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49300394</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49300394</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49300394</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>> The establishment Democrats are not complicit in Trump's nonsense<p>Oh yes they are! They vote to approve his cabinet picks, they campaign harder against left and progressive primary challengers than against the other party, they engage in the same insider trading as the other party, they take money from many of the same PACs and lobbying groups (Israel, AI, military-industrial complex), and...<p>> If you want to win in the general election, you want as many "normal" Democratic candidates as possible.<p>...and they continue to stick to the old strategy of "be centrist and try to flip upper-class urban moderate conservatives". This has been a proven consistent failure for 10 years now. Of course one can't expect every progressive and DSA leftist to win in November. But it's also not the 90s anymore. People still want to drain the swamp, as it were, and they don't see any hope in the establishment. They thought Trump was their anti-establishment man, but now, finally, after 10 years, they see what he really is, and there's an opportunity to <i>really</i> give the people what they want. And the best you can offer is more proven-failed centrism? Hell no. Schumer and Pelosi and Wasserman-Schultz and all the rest are complicit in the existential threat we face today. Primary them all out and send them to the dung heap of history.</p>
]]></description><pubDate>Fri, 14 Aug 2026 02:02:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49293987</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49293987</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49293987</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>It's an open question. Will Europe rally behind Denmark in WW3 NATO vs USA, or will we see transatlantic appeasement?</p>
]]></description><pubDate>Tue, 11 Aug 2026 17:46:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49261831</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49261831</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49261831</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>You can elect a <i>Congress</i> that will stand up to him in 2026. Step 1: continue primarying out establishment candidates from the Democratic Party, because they're largely complicit. Step 2: Vote for said non-establishment candidates in the general election.<p>The only challenge will be that the White House and complicit state governments are moving to disenfranchise as many people as possible before then, and likewise the White House will almost surely claim election fraud if Congress flips, and try to invalidate or even interfere with vote-counting.</p>
]]></description><pubDate>Tue, 11 Aug 2026 17:44:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49261805</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49261805</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49261805</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>Based on Google Translate it looks very much like what the article reports: Greenland Energy is moving exploratory drilling equipment without permits and without permission. Then it provides context on what Greenland Energy is and what they're up to.</p>
]]></description><pubDate>Tue, 11 Aug 2026 17:42:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49261770</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49261770</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49261770</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>And now in 2026 with NSPM-7 anyone who speaks up against literally anything the US does, can get put on a domestic terrorist list, which sets them up to be punished extrajudicially such as being de-banked. And then if the administration gets their way such individuals will be stripped of citizenship, detained in a concentration camp for a while because contractors make money off of that, and then eventually deported to some unknown place.</p>
]]></description><pubDate>Tue, 11 Aug 2026 17:38:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49261708</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49261708</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49261708</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>One correction: the US is not being run by an autocratic rapist. The US is being run by a cabal of authoritarian/monarchist billionaires, and the aforesaid autocratic rapist is just their representative in the executive branch, supported by the majority in Congress and the Supreme Court.</p>
]]></description><pubDate>Tue, 11 Aug 2026 17:34:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49261659</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49261659</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49261659</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>No news sites change headlines and edit articles all the time. Often there's a "last updated" date on the article and it can be days or weeks later than the original posting date. The last-updated date is better than nothing, but without a changelog it's very hard to know what version of the article you or anyone else read at any given time.</p>
]]></description><pubDate>Tue, 11 Aug 2026 14:47:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49259299</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49259299</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49259299</guid></item><item><title><![CDATA[New comment by gwerbin in "Greenland issues "strong warning" as American oil firm illegally drills wells"]]></title><description><![CDATA[
<p>While you're waffling on whether it's technically "illegal drilling" or "illegal not-yet-drilling":<p>1. the US president has previously threatened a war of conquest in Greenland, a territory of a NATO nation<p>2. the executive chair of the firm that is doing the illegal not-yet-drilling -- and also a top stakeholder -- is friends with said US president and donated to his campaigns.</p>
]]></description><pubDate>Tue, 11 Aug 2026 14:46:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49259285</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49259285</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49259285</guid></item><item><title><![CDATA[New comment by gwerbin in "US President hid in catering cart for secret flight in Turkey amid Iran threat"]]></title><description><![CDATA[
<p>I assume it's "as in travel". Sounds like the usual Trump ego thing. "It's dangerous, but I know you have to follow me because that's your job to report on what I'm doing, so now you're in danger."</p>
]]></description><pubDate>Tue, 11 Aug 2026 03:30:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49253034</link><dc:creator>gwerbin</dc:creator><comments>https://news.ycombinator.com/item?id=49253034</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49253034</guid></item></channel></rss>