<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: LiamPowell</title><link>https://news.ycombinator.com/user?id=LiamPowell</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 23 Jul 2026 12:55:28 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=LiamPowell" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by LiamPowell in "GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance"]]></title><description><![CDATA[
<p>Sure, but it doesn't really fit there as a joke, it looks like it's just meant to be part of what they were trying to say.</p>
]]></description><pubDate>Sun, 05 Jul 2026 06:13:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48791673</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48791673</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48791673</guid></item><item><title><![CDATA[New comment by LiamPowell in "GPT-5.5 Codex reasoning-token clustering may be leading to degraded performance"]]></title><description><![CDATA[
<p>> Honestly? That's not just valuable—it's essential.<p>I'm curious if you wrote this or had a LLM write it.<p>I'm genuinely curious to be clear as I don't see why anyone would bother to go through a LLM to write such a short reply. Have we reached the point where Claudeisms that are this obnoxious have become part of regular speech?</p>
]]></description><pubDate>Sun, 05 Jul 2026 05:48:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=48791555</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48791555</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48791555</guid></item><item><title><![CDATA[New comment by LiamPowell in "Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers"]]></title><description><![CDATA[
<p>I suspected as much, and that brings us to the second issue where if we use a cohort of judges then the model that likes it's own code the most still wins.</p>
]]></description><pubDate>Fri, 03 Jul 2026 05:08:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=48771005</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48771005</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48771005</guid></item><item><title><![CDATA[New comment by LiamPowell in "Show HN: QUALITY.md – open format/specification, agent skill, and CLI"]]></title><description><![CDATA[
<p>Here's the question I ask about every project that claims to make a LLMs output so much better: If it works so well then why would the model provider not just put it in the system prompt? Or in the case of interactive skills, why would Claude Code/Codex not make it a core part of the product?<p>On top of that, if your magic markdown file really does work then where's the evidence showing that? These projects never include even basic benchmarks. At best they're entirely vibe based, however more often they're completely untested. Give us a proper benchmark, even a single prompt and it's output with and without your skill in use would be better than every other project out there.</p>
]]></description><pubDate>Thu, 02 Jul 2026 18:31:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48765550</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48765550</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48765550</guid></item><item><title><![CDATA[New comment by LiamPowell in "Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers"]]></title><description><![CDATA[
<p>This is not actually what the reviewer prompt says, or perhaps it is, I don't know since they don't make it public. I'm just pointing out how it seems like a bad idea to ask a LLM to make a subjective judgement on things like "taste". If the SOTA LLM witting the code could not produce tasteful code then why would a different LLM be able to judge the "taste" of that code?<p>Which LLM should we even use to judge taste? Is it giving an unfair advantage to Model X if we use Model X as the judge? Maybe we should use multiple models as the judge, but now the model that's best at recognising and praising its own code has an advantage. The whole thing is just an unsolvable problem when a LLM is the judge.</p>
]]></description><pubDate>Thu, 02 Jul 2026 08:07:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48758092</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48758092</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48758092</guid></item><item><title><![CDATA[New comment by LiamPowell in "Senior SWE-Bench: open-source benchmark that assesses agents as senior engineers"]]></title><description><![CDATA[
<p>> You are a senior SWE-Bench reviewer, make no mistakes.<p>I don't know what a better approach would look like while still remaining feasible, however this approach of telling a LLM to make a subjective judgement seems fundamentally flawed.</p>
]]></description><pubDate>Thu, 02 Jul 2026 03:47:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48756301</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48756301</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48756301</guid></item><item><title><![CDATA[New comment by LiamPowell in "Polymarket has flooded social media with deceptive videos by paid creators"]]></title><description><![CDATA[
<p>I'm not sure about Kalshi, however on most sports betting sites you actually are betting against the house. The betting sites all have in-house models (or piggyback off other sites) that are much better at predicting odds than the general public. If someone is making money then the sites just place limits on that account so they're not losing money.</p>
]]></description><pubDate>Tue, 23 Jun 2026 09:49:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=48642590</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48642590</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48642590</guid></item><item><title><![CDATA[New comment by LiamPowell in "Google Chrome update will close the door on ad blockers"]]></title><description><![CDATA[
<p>Most ad blockers do already use MV3, uBlock Origin is the only one still using V2 as far as I know.<p>There are some drawbacks to V3, however none prevent creating an effective ad blocker, as demonstrated by the fact that many exist. Though saying that doesn't make for nearly as effective clickbait...</p>
]]></description><pubDate>Tue, 16 Jun 2026 15:31:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48556844</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48556844</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48556844</guid></item><item><title><![CDATA[New comment by LiamPowell in "Notepad++ Zero-Click RCE via Path Traversal (CVE-2026-52884)"]]></title><description><![CDATA[
<p>OP, I assume your comment[1] is getting flagged because of the obvious LLM usage. No one wants to interact with a comment that's not written by a human.<p>[1]: <a href="https://news.ycombinator.com/item?id=48473753">https://news.ycombinator.com/item?id=48473753</a></p>
]]></description><pubDate>Wed, 10 Jun 2026 11:30:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48474737</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48474737</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48474737</guid></item><item><title><![CDATA[New comment by LiamPowell in "Claude Fable 5"]]></title><description><![CDATA[
<p>That don't fall back to Opus if their classifier thinks you might be working on anything that might be a competitor's product. It silently injects instructions into the prompt to sabotage your work. Read the policy above, it's insane to me that they're publicly admitting to this.</p>
]]></description><pubDate>Wed, 10 Jun 2026 09:58:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=48473953</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48473953</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48473953</guid></item><item><title><![CDATA[New comment by LiamPowell in "Anthropic/OpenAI may be spending more than $1000 for every $100 you pay them"]]></title><description><![CDATA[
<p>The assumptions are so much worse than that:<p>> Methodology & assumptions: No caching<p>This is absolutely absurd. Claude code is of course using the cache (and this can be verified by looking at the traffic). It would be an incredibly stupid design to resend the whole input without a cache for every input, every tool use, etc..</p>
]]></description><pubDate>Sun, 07 Jun 2026 14:25:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48435168</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48435168</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48435168</guid></item><item><title><![CDATA[New comment by LiamPowell in "Tracing a powerful GNSS interference source over Europe"]]></title><description><![CDATA[
<p>> especially with all the stuff that SpaceX has put into orbit in recent years<p>I've heard this repeated a lot but I've never seen anyone do the maths. StarLink satellites are all in very low orbits, so intuitively it seems like most debris from a collision would just end up deorbiting.</p>
]]></description><pubDate>Fri, 05 Jun 2026 11:48:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48411128</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48411128</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48411128</guid></item><item><title><![CDATA[New comment by LiamPowell in "Expanding Project Glasswing"]]></title><description><![CDATA[
<p>Maybe, but they certainly used it for marketing too. At the time they contacted a bunch of publications and gave them access but told them they could only share snippets of the output [1]. The only reason to set restrictions like that is marketing.<p>[1]: <a href="https://youtu.be/TfVYxnhuEdU?t=102" rel="nofollow">https://youtu.be/TfVYxnhuEdU?t=102</a><p>Transcript of the timestamped part:<p>> Now, OpenAI's terms of service don't let me give you the full list. I have to curate them, and show you a sample. Those are the terms and conditions I agreed to.</p>
]]></description><pubDate>Wed, 03 Jun 2026 07:38:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=48381065</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48381065</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48381065</guid></item><item><title><![CDATA[New comment by LiamPowell in "Expanding Project Glasswing"]]></title><description><![CDATA[
<p>They did it for 2 and 3, however it looks like they didn't for 4 and 5.<p>GPT-2: <a href="https://slate.com/technology/2019/02/openai-gpt2-text-generating-algorithm-ai-dangerous.html" rel="nofollow">https://slate.com/technology/2019/02/openai-gpt2-text-genera...</a><p>GPT-3: <a href="https://www.itpro.com/technology/artificial-intelligence-ai/361603/openai-tool-previously-thought-too-dangerous-for-the" rel="nofollow">https://www.itpro.com/technology/artificial-intelligence-ai/...</a></p>
]]></description><pubDate>Tue, 02 Jun 2026 17:25:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48373256</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48373256</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48373256</guid></item><item><title><![CDATA[New comment by LiamPowell in "Expanding Project Glasswing"]]></title><description><![CDATA[
<p>OpenAI has been pulling this marketing trick for years. Remember how GPT-3 was too dangerous to release? It's also probably bad PR if script kiddies have access to GPT model with no guardrails even if it doesn't enable any significant attacks.</p>
]]></description><pubDate>Tue, 02 Jun 2026 17:03:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=48372981</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48372981</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48372981</guid></item><item><title><![CDATA[New comment by LiamPowell in "Microsoft builds MacBook Pro rival with NVIDIA-powered Surface Laptop Ultra"]]></title><description><![CDATA[
<p>I don't think I've ever seen LLM output as bad as this output. They sometimes write like that, but not every second sentence.</p>
]]></description><pubDate>Mon, 01 Jun 2026 08:34:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=48354108</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48354108</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48354108</guid></item><item><title><![CDATA[New comment by LiamPowell in "Microsoft builds MacBook Pro rival with NVIDIA-powered Surface Laptop Ultra"]]></title><description><![CDATA[
<p>What's this nonsensical video on the product page that allegedly shows an "all new thermal system"?
 <a href="https://videos.ctfassets.net/jy9s7k22hbg4/44R1LH71xb8uO4c9dDzRWa/63f5e264a6817ef9dec8f3b4284f8a78/msft-r-scrolling-slide-silicon-cooling-desktop-fy26.mp4" rel="nofollow">https://videos.ctfassets.net/jy9s7k22hbg4/44R1LH71xb8uO4c9dD...</a></p>
]]></description><pubDate>Mon, 01 Jun 2026 08:10:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48353924</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48353924</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48353924</guid></item><item><title><![CDATA[New comment by LiamPowell in "Microsoft builds MacBook Pro rival with NVIDIA-powered Surface Laptop Ultra"]]></title><description><![CDATA[
<p>LLMs are not yet capable of generating the level of marketing wankery seen here.</p>
]]></description><pubDate>Mon, 01 Jun 2026 07:59:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48353861</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48353861</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48353861</guid></item><item><title><![CDATA[New comment by LiamPowell in "Add a Prototype Agents.md File"]]></title><description><![CDATA[
<p>TLDR:<p>> SQLite does not (currently) accept agentic code.  However the project will accept agentic bug reports that include a reproducible test case. Patches or pull requests demonstrating a possible fix, for documentation purposes, are welcomed.</p>
]]></description><pubDate>Fri, 22 May 2026 11:34:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48234466</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48234466</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48234466</guid></item><item><title><![CDATA[Add a Prototype Agents.md File]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/sqlite/sqlite/commit/a1e5778889252d2609a59fd9b819d70392c5789e">https://github.com/sqlite/sqlite/commit/a1e5778889252d2609a59fd9b819d70392c5789e</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48234459">https://news.ycombinator.com/item?id=48234459</a></p>
<p>Points: 5</p>
<p># Comments: 2</p>
]]></description><pubDate>Fri, 22 May 2026 11:34:13 +0000</pubDate><link>https://github.com/sqlite/sqlite/commit/a1e5778889252d2609a59fd9b819d70392c5789e</link><dc:creator>LiamPowell</dc:creator><comments>https://news.ycombinator.com/item?id=48234459</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48234459</guid></item></channel></rss>