<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: km144</title><link>https://news.ycombinator.com/user?id=km144</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 24 Sep 2026 01:13:02 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=km144" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by km144 in "Claude Opus 5.5"]]></title><description><![CDATA[
<p>I think this release is really going to give them a hard time selling Fable:<p>> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.<p>In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.</p>
]]></description><pubDate>Tue, 22 Sep 2026 16:41:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49804122</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49804122</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49804122</guid></item><item><title><![CDATA[New comment by km144 in "Claude Opus 5.5"]]></title><description><![CDATA[
<p>Can you fix the link on <i>that</i> post then? I duped because that post links to a diff that tells me nothing about Opus 5.5</p>
]]></description><pubDate>Tue, 22 Sep 2026 16:35:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49803995</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49803995</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49803995</guid></item><item><title><![CDATA[Claude Opus 5.5]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.anthropic.com/claude-opus-5-5">https://www.anthropic.com/claude-opus-5-5</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49803892">https://news.ycombinator.com/item?id=49803892</a></p>
<p>Points: 1768</p>
<p># Comments: 1081</p>
]]></description><pubDate>Tue, 22 Sep 2026 16:29:05 +0000</pubDate><link>https://www.anthropic.com/claude-opus-5-5</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49803892</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49803892</guid></item><item><title><![CDATA[New comment by km144 in "A coffee shop owner used AI to make a menu poster. Then came the angry DMs"]]></title><description><![CDATA[
<p>When I read that, the only thing I could think was "holy shit get off your fucking phone". But I agree with you as well.</p>
]]></description><pubDate>Thu, 17 Sep 2026 12:25:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49739767</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49739767</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49739767</guid></item><item><title><![CDATA[New comment by km144 in "Astra for Coding: Why Are We Doing This Again?"]]></title><description><![CDATA[
<p>The biggest jump was Opus 4.6. Since then they have gradually gotten better at finding issues in your reasoning, not hallucinating, and being rigorous with the code, but much much worse at explaining things and generally just talking in a way that a human can understand. All the models I've tried seem to be suffering from the same fate so it must be something going on with the training meta right now.</p>
]]></description><pubDate>Fri, 11 Sep 2026 19:45:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49664299</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49664299</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49664299</guid></item><item><title><![CDATA[New comment by km144 in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>Yep. Simple answer is they want to IPO in the fall, and a new Haiku does literally nothing for them</p>
]]></description><pubDate>Tue, 01 Sep 2026 18:37:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49526102</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49526102</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49526102</guid></item><item><title><![CDATA[New comment by km144 in "Fastpotify"]]></title><description><![CDATA[
<p>That would actually be a reasonable thing to do with a model that is good at writing though. I'm sure if you used that method on Opus 4.6 you'd get a pretty decent result, because Claude models used to be quite good at that sort of technical synthesis. Asking a model that is bad at technical writing (e.g. Opus/Fable 5) to write your docs is a bad idea.</p>
]]></description><pubDate>Tue, 01 Sep 2026 14:14:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49522314</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49522314</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49522314</guid></item><item><title><![CDATA[New comment by km144 in "Fastpotify"]]></title><description><![CDATA[
<p>I believe AI technical prose peaked with Opus 4.6. I still use it (mostly for that purpose) and I think it's legitimately great. I'm hoping they are able to reverse the trends that benchmaxxing and RLAIF have wrought on Claude's non-code output with the future models.</p>
]]></description><pubDate>Tue, 01 Sep 2026 14:11:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49522261</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49522261</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49522261</guid></item><item><title><![CDATA[New comment by km144 in "OpenLogi"]]></title><description><![CDATA[
<p>I assure you it won't be as bad as this (answer to FAQ: "Will OpenLogi support Logitech Flow?")<p>> It's on the roadmap, at the far end: a cross-computer pointer and clipboard bridge is a very large feature. The half that lives in the protocol already ships. OpenLogi drives Easy-Switch host switching over HID++ (0x1814/0x1815), and paired mice follow the keyboard when it switches hosts. If the rest lands, it will be opt-in and local-network only.<p>A great example of why Claude's writing is terrible is the sentence "The half that lives in the protocol already ships". I cannot imagine any human writing this.</p>
]]></description><pubDate>Wed, 19 Aug 2026 16:36:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49363809</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49363809</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49363809</guid></item><item><title><![CDATA[New comment by km144 in "Qwen3.8 Max now ranked as the best overall model by agentic index"]]></title><description><![CDATA[
<p>The entire ecosystem of CC is designed to facilitate burning tokens. You have to ask the LLM to write a script for the app to tell you which folder you're working in and which branch you're on. There are commands that just diagnose your Claude Code setup and try to "optimize" it. Adding skills or plugins bloats the context window. Developing plans means that you work through questions before you get to it in the code, but that matters way more for human programmers than LLMs, so it's probably just a waste of tokens.</p>
]]></description><pubDate>Fri, 07 Aug 2026 14:33:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49211081</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49211081</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49211081</guid></item><item><title><![CDATA[New comment by km144 in "Qwen3.8 Max now ranked as the best overall model by agentic index"]]></title><description><![CDATA[
<p>What output style have you found to actually fix Opus 5's grating prose then? I find it leans hard into its preferred grammatical structures and rote sayings no matter what I include in the output style.</p>
]]></description><pubDate>Fri, 07 Aug 2026 14:30:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49211025</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49211025</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49211025</guid></item><item><title><![CDATA[New comment by km144 in "Not hiring junior engineers won't solve the problem you think you have"]]></title><description><![CDATA[
<p>> because it's really hard to learn anything by skimming code that's being pooped out by Claude.<p>I think this is pretty dubious. It's true that you have to do more proactive learning with an agentic workflow, but it doesn't mean you can't learn stuff. Software engineers forged in the age of AI are not learning the same things their predecessors did or thinking about code in the same way, but the more curious people will still be better at the job, which is I think how it's always been. It seems a little unclear how much "AI brain rot" will be a force unto itself that neuters even the smartest of the juniors, but I think it's still an open question.</p>
]]></description><pubDate>Wed, 05 Aug 2026 19:37:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49187850</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49187850</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49187850</guid></item><item><title><![CDATA[New comment by km144 in "Claude: Elevated errors across all models – Resolved"]]></title><description><![CDATA[
<p>You're right—and it's worth calling out explicitly, because that completely changes the approach here.</p>
]]></description><pubDate>Wed, 29 Jul 2026 20:21:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49102527</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49102527</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49102527</guid></item><item><title><![CDATA[New comment by km144 in "Claude Opus 5"]]></title><description><![CDATA[
<p>Here is one data point for cost:<p><a href="https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task&model-filters=large-models%2Cproprietary%2Creasoning-models#cost-tabs" rel="nofollow">https://artificialanalysis.ai/models?cost=intelligence-vs-co...</a><p>Here is another data point for output token efficiency:<p><a href="https://artificialanalysis.ai/models?cost=intelligence-vs-cost-per-task&model-filters=large-models%2Cproprietary%2Creasoning-models&intelligence-index-token-use=intelligence-vs-output-tokens-per-task#intelligence-index-token-use-tabs" rel="nofollow">https://artificialanalysis.ai/models?cost=intelligence-vs-co...</a></p>
]]></description><pubDate>Fri, 24 Jul 2026 18:04:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49039456</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49039456</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49039456</guid></item><item><title><![CDATA[New comment by km144 in "Claude Opus 5"]]></title><description><![CDATA[
<p>I agree, very odd they did not comment on any theories for the degradation here. Dip and then rebound at max effort is pretty interesting too. Overthinking is bad, but you can overthink so much it starts to be better again?</p>
]]></description><pubDate>Fri, 24 Jul 2026 17:59:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49039400</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=49039400</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49039400</guid></item><item><title><![CDATA[New comment by km144 in "Kernel anti-cheat is an overreach"]]></title><description><![CDATA[
<p>> So I would rather share a match with the occasional cheater than run un-auditable ring-0 software on the same machine I use for anything private.<p>The article makes an argument that anti-cheat is not worth the trade-off, yet the author admits they are a non-gamer. Then they go on to present one example of anti-cheat that tells us all we need to know about actual gamers' preferences—FACEIT. For those who don't know, FACEIT is a third-party matchmaking service, primarily for CS2. People <i>choose</i> to go through the hoops of using third-party service that installs kernel-level anti-cheat on their computer because it helps to keep cheaters out of their games. This seems like pretty strong evidence that the author's argument is not a good representation of gamers' thoughts on this. I don't know what the actual solution is. I suspect if Valve made their own kernel-level anti-cheat people might trust it more, but it's still the same problem.</p>
]]></description><pubDate>Tue, 07 Jul 2026 12:05:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48816579</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=48816579</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48816579</guid></item><item><title><![CDATA[New comment by km144 in "Artificial intelligence is not conscious – Ted Chiang"]]></title><description><![CDATA[
<p>I'm basically restating what you said, but it's amazing to me that the vast majority of people you will meet, even educated people, are casual dualists and free-will libertarians. If they happen to acknowledge materialism in some way (i.e. the acceptance of the idea that the brain's processes are just the interaction of physical matter), there is still zero chance they draw determinist conclusions from that acknowledgement. But I guess that tracks, given that most professional philosophers are apparently compatibilists for some reason I have never understood (the arguments get really confusing).</p>
]]></description><pubDate>Thu, 04 Jun 2026 17:32:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48401862</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=48401862</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48401862</guid></item><item><title><![CDATA[New comment by km144 in "Can we have the day off?"]]></title><description><![CDATA[
<p>Because, as the OP said, working hours are powered by norms. There are salaried positions and companies and teams that certainly will make you work 6 days a week, or make you feel like a bad worker if you don't do anything on a Saturday. But the vast majority of companies (and employees within those companies) would consider the expectation of working a 6th day to be completely unacceptable.</p>
]]></description><pubDate>Thu, 28 May 2026 16:34:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48311396</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=48311396</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48311396</guid></item><item><title><![CDATA[New comment by km144 in "Gemini 3.1 Pro"]]></title><description><![CDATA[
<p>The real question is: Why are people designing benchmarks that, if a model is trained on them, it won't improve the performance of the model at any real-world tasks? Why would anyone care about such benchmarks?</p>
]]></description><pubDate>Fri, 20 Feb 2026 16:09:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=47089857</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=47089857</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47089857</guid></item><item><title><![CDATA[New comment by km144 in "Court orders restart of all US offshore wind power construction"]]></title><description><![CDATA[
<p>But there is a way for even an aligned federal government to fight back against the slide into authoritarianism, even with an authoritarian president expanding the powers of the executive, and that is for the other branches to strongly advocate for their own power. The problem as I see it is that Congress literally does not care that they are ceding more power than ever before to the executive. Mostly I think this is due to the cult of personality aspect of Trumpism and the idea that you're basically either with him and in the party or against him and out of the party, so it's impossible to drum up support within the party to fight back against the wresting of power. But also it's because the Republican party has no interest in actually passing legislation because most non-budgetary directions they can go will result in incredible cross-pressure (healthcare reform, federal abortion bans, etc). They believe they are better off not doing policy and letting Trump do whatever.</p>
]]></description><pubDate>Tue, 03 Feb 2026 15:56:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=46872616</link><dc:creator>km144</dc:creator><comments>https://news.ycombinator.com/item?id=46872616</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46872616</guid></item></channel></rss>