<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: zone411</title><link>https://news.ycombinator.com/user?id=zone411</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 05 Sep 2026 09:22:57 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=zone411" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by zone411 in "OpenAI's GPT-6 Astra on ARC-AGI-3"]]></title><description><![CDATA[
<p>I don't:<p>"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.<p>Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.</p>
]]></description><pubDate>Fri, 04 Sep 2026 03:23:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49560161</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=49560161</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49560161</guid></item><item><title><![CDATA[New comment by zone411 in "Intelligence Is Not the Main Bottleneck"]]></title><description><![CDATA[
<p>A complete mischaracterization, as usual for HN lately when discussing AI or LessWrong. Obviously, even average levels of persuasion are enough to convince some people. And nobody is air-gapping AI.</p>
]]></description><pubDate>Wed, 05 Aug 2026 16:16:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49184921</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=49184921</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49184921</guid></item><item><title><![CDATA[New comment by zone411 in "Be skeptical of OpenAI's rogue hacker agent story"]]></title><description><![CDATA[
<p>So why don't companies in other industries rush to prove their products are dangerous weapons? Maybe because it would be a really dumb PR stunt?</p>
]]></description><pubDate>Sat, 25 Jul 2026 04:46:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49044612</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=49044612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49044612</guid></item><item><title><![CDATA[New comment by zone411 in "Be skeptical of OpenAI's rogue hacker agent story"]]></title><description><![CDATA[
<p>And how do people saying this know the capabilities of yet unreleased models?</p>
]]></description><pubDate>Sat, 25 Jul 2026 04:44:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49044610</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=49044610</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49044610</guid></item><item><title><![CDATA[New comment by zone411 in "Be skeptical of OpenAI's rogue hacker agent story"]]></title><description><![CDATA[
<p>How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?</p>
]]></description><pubDate>Sat, 25 Jul 2026 04:43:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49044603</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=49044603</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49044603</guid></item><item><title><![CDATA[New comment by zone411 in "Be skeptical of OpenAI's rogue hacker agent story"]]></title><description><![CDATA[
<p>It's not good marketing for them. This is a talking point with zero evidence that people repeat mindlessly. Scaring customers, worrying employees, and inviting regulators to act would be the worst marketing idea ever devised.</p>
]]></description><pubDate>Sat, 25 Jul 2026 04:41:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49044587</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=49044587</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49044587</guid></item><item><title><![CDATA[Natural-Density Almost-Bounded Collatz Orbits in Logarithmic Time (AI, Lean)]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.proofatlas.ai/formalizations/natural-density-log-time-collatz/">https://www.proofatlas.ai/formalizations/natural-density-log-time-collatz/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48996871">https://news.ycombinator.com/item?id=48996871</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 21 Jul 2026 19:18:43 +0000</pubDate><link>https://www.proofatlas.ai/formalizations/natural-density-log-time-collatz/</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=48996871</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48996871</guid></item><item><title><![CDATA[New comment by zone411 in "Openrouter Fusion API"]]></title><description><![CDATA[
<p>Yes, definitely not a new idea. I had a multi-turn composite model in 2024 that was outperforming the top models across benchmarks: <a href="https://x.com/LechMazur/status/1828804485033992514" rel="nofollow">https://x.com/LechMazur/status/1828804485033992514</a>.</p>
]]></description><pubDate>Mon, 15 Jun 2026 17:16:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48544298</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=48544298</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48544298</guid></item><item><title><![CDATA[New comment by zone411 in "Artificial intelligence is not conscious – Ted Chiang"]]></title><description><![CDATA[
<p>That's not proof. Emergent intelligence is not consciousness.</p>
]]></description><pubDate>Thu, 04 Jun 2026 17:16:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48401633</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=48401633</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48401633</guid></item><item><title><![CDATA[New comment by zone411 in "The mysterious Hy3 LLM is topping OpenRouter Model Rankings by a large margin"]]></title><description><![CDATA[
<p>I’ve tested this model on four of my benchmarks:<p><a href="https://github.com/lechmazur/buyout_game" rel="nofollow">https://github.com/lechmazur/buyout_game</a> 10th out 36.<p><a href="https://github.com/lechmazur/pact/" rel="nofollow">https://github.com/lechmazur/pact/</a> 14th out 25.<p><a href="https://github.com/lechmazur/nyt-connections/" rel="nofollow">https://github.com/lechmazur/nyt-connections/</a> 60th out 81.<p><a href="https://github.com/lechmazur/debate" rel="nofollow">https://github.com/lechmazur/debate</a> 16th out of 29.</p>
]]></description><pubDate>Fri, 29 May 2026 05:10:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48319302</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=48319302</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48319302</guid></item><item><title><![CDATA[New comment by zone411 in "Spain blocks prediction markets Polymarket, Kalshi over lack of gambling licence"]]></title><description><![CDATA[
<p>100%. It's sad to see that this attitude has spread to HN</p>
]]></description><pubDate>Tue, 26 May 2026 18:22:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=48283683</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=48283683</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48283683</guid></item><item><title><![CDATA[New comment by zone411 in "An OpenAI model has disproved a central conjecture in discrete geometry"]]></title><description><![CDATA[
<p>I actually tried using GPT-5.5 Pro on this problem recently. It thought it was making progress on one path, but it made so many mistakes that it didn't feel worth it pushing further. It'll be interesting to check whether it's the same route. I got partial results (proved in Lean) that improve on the best-known results for four Erdős problems with GPT-5.5 Pro</p>
]]></description><pubDate>Wed, 20 May 2026 23:18:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48215681</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=48215681</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48215681</guid></item><item><title><![CDATA[LLM Position Bias Benchmark: Swapped-Order Pairwise Judging]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/lechmazur/position_bias">https://github.com/lechmazur/position_bias</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47855226">https://news.ycombinator.com/item?id=47855226</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 21 Apr 2026 22:08:12 +0000</pubDate><link>https://github.com/lechmazur/position_bias</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47855226</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47855226</guid></item><item><title><![CDATA[New comment by zone411 in "EFF is leaving X"]]></title><description><![CDATA[
<p><a href="https://variety.com/2020/digital/news/twitter-unblocks-new-york-post-hunter-biden-hacked-materials-1234820449/" rel="nofollow">https://variety.com/2020/digital/news/twitter-unblocks-new-y...</a></p>
]]></description><pubDate>Thu, 09 Apr 2026 18:25:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=47707539</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47707539</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47707539</guid></item><item><title><![CDATA[Show HN: Buyout Game Benchmark: Multi-Agent Bargaining, Transfers, and Takeovers]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/lechmazur/buyout_game">https://github.com/lechmazur/buyout_game</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47580319">https://news.ycombinator.com/item?id=47580319</a></p>
<p>Points: 6</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 30 Mar 2026 22:07:13 +0000</pubDate><link>https://github.com/lechmazur/buyout_game</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47580319</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47580319</guid></item><item><title><![CDATA[New comment by zone411 in "AI overly affirms users asking for personal advice"]]></title><description><![CDATA[
<p>I built this benchmark this month: <a href="https://github.com/lechmazur/sycophancy" rel="nofollow">https://github.com/lechmazur/sycophancy</a>. There are large differences between LLMs. There are large differences between LLMs. For example, Mistral Large 3 and GPT-4.1 will initially agree with the narrator, while Gemini will disagree. I swap sides, so this is not about possible viewpoint bias in the LLMs. But another benchmark shows that Gemini will then change its view very easily in a multi-turn conversation while Kimi K2.5 or Grok won't: <a href="https://github.com/lechmazur/persuasion" rel="nofollow">https://github.com/lechmazur/persuasion</a>.</p>
]]></description><pubDate>Sat, 28 Mar 2026 16:51:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=47556280</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47556280</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47556280</guid></item><item><title><![CDATA[New comment by zone411 in "Folk are getting dangerously attached to AI that always tells them they're right"]]></title><description><![CDATA[
<p>I built two related benchmarks this month: <a href="https://github.com/lechmazur/sycophancy" rel="nofollow">https://github.com/lechmazur/sycophancy</a> and <a href="https://github.com/lechmazur/persuasion" rel="nofollow">https://github.com/lechmazur/persuasion</a>. There are large differences between LLMs. For example, good luck getting Grok to change its view, while Gemini 3.1 Pro will usually disagree with the narrator at first but then change its position very easily when pushed.</p>
]]></description><pubDate>Sat, 28 Mar 2026 16:45:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=47556226</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47556226</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47556226</guid></item><item><title><![CDATA[LLM Persuasion Benchmark: Multi-Turn Persuasion Between Models]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/lechmazur/persuasion">https://github.com/lechmazur/persuasion</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47545308">https://news.ycombinator.com/item?id=47545308</a></p>
<p>Points: 9</p>
<p># Comments: 0</p>
]]></description><pubDate>Fri, 27 Mar 2026 17:02:35 +0000</pubDate><link>https://github.com/lechmazur/persuasion</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47545308</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47545308</guid></item><item><title><![CDATA[New comment by zone411 in "Show HN: LLM Debate Benchmark"]]></title><description><![CDATA[
<p>Hmm, maybe in the next edition, Opus gets expensive. I should probably run GPT-5.4 xhigh too if I do that for fairness...</p>
]]></description><pubDate>Mon, 23 Mar 2026 21:10:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=47495156</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47495156</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47495156</guid></item><item><title><![CDATA[Show HN: LLM Debate Benchmark]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/lechmazur/debate/">https://github.com/lechmazur/debate/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47494895">https://news.ycombinator.com/item?id=47494895</a></p>
<p>Points: 9</p>
<p># Comments: 3</p>
]]></description><pubDate>Mon, 23 Mar 2026 20:49:45 +0000</pubDate><link>https://github.com/lechmazur/debate/</link><dc:creator>zone411</dc:creator><comments>https://news.ycombinator.com/item?id=47494895</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47494895</guid></item></channel></rss>