<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: buttered_toast</title><link>https://news.ycombinator.com/user?id=buttered_toast</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 04 Aug 2026 12:23:47 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=buttered_toast" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by buttered_toast in "Gemini 3.1 Pro"]]></title><description><![CDATA[
<p>Makes me wonder what people would consider better, a model that gets 92% of questions right 100% of the time, or a model that gets 95% of the questions right 90% of the time and 88% right the other 10%?<p>I think that's why benchmarking is so hard for me to fully get behind, even if we do it over say, 20 attempts and average it. For a given model, those 20 attempts could have had 5 incredible outcomes and 15 mediocre ones, whereas another model could have 20 consistently decent attempts and the average score would be generally the same.<p>We at least see variance in public benchmarks, but in the internal examples that's almost never the case.</p>
]]></description><pubDate>Thu, 19 Feb 2026 19:05:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=47077701</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47077701</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47077701</guid></item><item><title><![CDATA[New comment by buttered_toast in "Gemini 3.1 Pro"]]></title><description><![CDATA[
<p>I think we need to reevaluate what purpose these sorts of questions serve and why they're important in regards to judging intelligence.<p>The model getting it correct or not at any given instance isn't the point, the point is if the model ever gets it wrong we can still assume that it still has some semblance of stochasticity in its output, given that a model is essentially static once it is released.<p>Additionally, hey don't learn post training (except for in context which I think counts as learning to some degree albeit transient), if hypothetically it answers incorrectly 1 in 50 attempts, and I explain in that 1 failed attempt why it is wrong, it will still be a 1-50 chance it gets it wrong in a new instance.<p>This differs from humans, say for example I give an average person the "what do you put in a toaster" trick and they fall for it, I can be pretty confident that if I try that trick again 10 years later they will probably not fall for it, you can't really say that for a given model.</p>
]]></description><pubDate>Thu, 19 Feb 2026 18:45:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=47077418</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47077418</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47077418</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>I think it says in the paper that he does, but it's also public knowledge.<p><a href="https://www.linkedin.com/in/alex-lupsasca-9096a214/" rel="nofollow">https://www.linkedin.com/in/alex-lupsasca-9096a214/</a></p>
]]></description><pubDate>Fri, 13 Feb 2026 23:32:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=47009306</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47009306</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47009306</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>Okay I see what you mean, and yeah that sounds reasonable too. Do you have any context on that first part? I would like to know more about how/why they might not have been able to pursue more training runs.</p>
]]></description><pubDate>Fri, 13 Feb 2026 22:19:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=47008580</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47008580</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47008580</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>Thank you for taking the time to reply, I see you might have already answered this elsewhere so it's much appreciated.</p>
]]></description><pubDate>Fri, 13 Feb 2026 22:00:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=47008415</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47008415</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47008415</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>Oh that's really cool, I am not versed in physics by any means, can you explain how you believed there to be a simple formula but were unable to find it? What would lead you to believe that instead of just accepting it at face value?</p>
]]></description><pubDate>Fri, 13 Feb 2026 21:48:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=47008290</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47008290</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47008290</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>Can't say, just seems implausible, but I am a nobody anyways ¯\_(ツ)_/¯</p>
]]></description><pubDate>Fri, 13 Feb 2026 20:25:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=47007381</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47007381</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47007381</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>I would interpret it as implying that the result was due to a lot more hand-holding that what is let on.<p>Was the initial conjecture based on leading info from the other authors or was it simply the authors presenting all information and asking for a conjecture?<p>Did the authors know that there was a simpler means of expressing the conjecture and lead GPT to its conclusion, or did it spontaneously do so on its own after seeing the hand-written expressions.<p>These aren't my personal views, but there is some handwaving about the process in such a way that reads as if this was all spontaneous involvement on GPTs end.<p>But regardless, a result is a result so I'm content with it.</p>
]]></description><pubDate>Fri, 13 Feb 2026 20:23:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=47007359</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47007359</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47007359</guid></item><item><title><![CDATA[New comment by buttered_toast in "GPT-5.2 derives a new result in theoretical physics"]]></title><description><![CDATA[
<p>Absolutely no way this is true right? Ilya left around the time 4o was released. I can't imagine they haven't had a single successful run since then.</p>
]]></description><pubDate>Fri, 13 Feb 2026 20:02:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=47007085</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47007085</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47007085</guid></item><item><title><![CDATA[New comment by buttered_toast in "Gemini 3 Deep Think"]]></title><description><![CDATA[
<p>Thank you!</p>
]]></description><pubDate>Fri, 13 Feb 2026 16:40:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=47004667</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47004667</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47004667</guid></item><item><title><![CDATA[New comment by buttered_toast in "Gemini 3 Deep Think"]]></title><description><![CDATA[
<p>Couldn't you just make up new combinations, or new caveats indefinitely to mitigate that? It would be nice to see maybe 3-4 good examples for validation. I'd do it myself, but I don't have $200 to play around with this model.</p>
]]></description><pubDate>Fri, 13 Feb 2026 15:54:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=47004134</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47004134</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47004134</guid></item><item><title><![CDATA[New comment by buttered_toast in "Gemini 3 Deep Think"]]></title><description><![CDATA[
<p>Is there a way you can showcase a few of these?</p>
]]></description><pubDate>Fri, 13 Feb 2026 14:25:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=47003104</link><dc:creator>buttered_toast</dc:creator><comments>https://news.ycombinator.com/item?id=47003104</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47003104</guid></item></channel></rss>