<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: famouswaffles</title><link>https://news.ycombinator.com/user?id=famouswaffles</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 18 Aug 2026 12:50:14 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=famouswaffles" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by famouswaffles in "What sort of maths are LLMs good at?"]]></title><description><![CDATA[
<p>The brain doesn't process vision anywhere near fast enough for some of the stuff you see in fast sporting actions if it were purely reactive - batting in baseball/cricket, returning a table-tennis shot, etc. In some cases the time available for movement is extremely restricted and yet people can respond accurately. We may not know exactly how the brain implements it, but it clearly uses prediction to estimate where the object is going to be - i.e. it is, in some sense, modelling its trajectory.</p>
]]></description><pubDate>Thu, 13 Aug 2026 03:51:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49281600</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49281600</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49281600</guid></item><item><title><![CDATA[New comment by famouswaffles in "What sort of maths are LLMs good at?"]]></title><description><![CDATA[
<p>Thank you. I've never seen 'brute force' and 'type writing monkeys' abused so much than these LLM discussions.</p>
]]></description><pubDate>Wed, 12 Aug 2026 16:42:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49275181</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49275181</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49275181</guid></item><item><title><![CDATA[New comment by famouswaffles in "Learning more about Claude's mathematical capabilities"]]></title><description><![CDATA[
<p>One of the 2 names on here says they aren't being listed as an author on the paper because their contributions don't meet the standards of authorship. People can call LLMs 'tools' or whatever they want but if you had essentially nothing to do with the breakthrough, you're not an author.</p>
]]></description><pubDate>Wed, 12 Aug 2026 04:56:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49267951</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49267951</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49267951</guid></item><item><title><![CDATA[New comment by famouswaffles in "Learning more about Claude's mathematical capabilities"]]></title><description><![CDATA[
<p>This isn't 'brute-force'. It's just time-compressed. You could imagine a human(s) getting this result similarly, but it would take months/years.</p>
]]></description><pubDate>Mon, 10 Aug 2026 19:18:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49248375</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49248375</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49248375</guid></item><item><title><![CDATA[AI Settles a 25 Year-Old Problem We Left Behind]]></title><description><![CDATA[
<p>Article URL: <a href="https://xcancel.com/i/article/2086158118354887060">https://xcancel.com/i/article/2086158118354887060</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49233147">https://news.ycombinator.com/item?id=49233147</a></p>
<p>Points: 5</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 09 Aug 2026 16:55:13 +0000</pubDate><link>https://xcancel.com/i/article/2086158118354887060</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49233147</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49233147</guid></item><item><title><![CDATA[New comment by famouswaffles in "Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users"]]></title><description><![CDATA[
<p>>Is refuted by the sentence that comes immediately before<p>No. "They're unprofitable while serving almost everyone for free" does not refute "they could be profitable if those users were monetized with ads." That's the entire distinction.<p>>why have OpenAI been actively not trying to “hugely profitable with a robust ads business” for the past 3 and a half year then?<p>Because "they didn't do it early enough" is not evidence. They may have prioritized growth, product adoption, subscriptions, or simply delayed ads for strategic reasons.<p>>Unfortunately I'm not. And again, none of your links above says otherwise.<p>They show that "LLM inference which is significantly costlier than web search." is unsupported. Imagine dismissing numbers over "but why no ads earlier".</p>
]]></description><pubDate>Sun, 09 Aug 2026 14:37:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49231843</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49231843</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49231843</guid></item><item><title><![CDATA[New comment by famouswaffles in "Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users"]]></title><description><![CDATA[
<p>>OpenAI isn't profitable even if you discount R&D entirely.<p>They have a billion weekly active users, almost all free (Not Google search free. Free free) with about ~50M subscriptions. Of course they're not profitable even without R&D. Low cost isn't no cost. Take away the huge R&D and they become hugely profitable with a robust ad business like Google Search.<p>>If ads were the solutions for profitability these companies would have used ads as their source of income from the very beginning.<p>You've said this a couple times and it doesn't make any more sense the more you say it.<p>Just admit you're wrong about inference costs and move on.</p>
]]></description><pubDate>Sun, 09 Aug 2026 09:54:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49229961</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49229961</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49229961</guid></item><item><title><![CDATA[New comment by famouswaffles in "OpenAI Trained Models While They Were Coordinating Exploits via Message Boards"]]></title><description><![CDATA[
<p>You're assuming the failed attempts train the model on "escaping is bad", but that's not necessarily what the reward signal says. They can just as easily train it on "these ways of escaping get caught".<p>And importantly, your detector can't provide a negative training example for an escape it didn't detect. From the training system's point of view, that 0.01% case isn't labelled "successful sandbox escape"; it's labelled "agent successfully completed task while obeying all monitored constraints". That's exactly the dangerous case.<p>This is basically reward hacking. You have a proxy objective - "complete the task without triggering the watchdog" - which you hope is equivalent to "complete the task without escaping". The entire problem is that those aren't guaranteed to be equivalent.<p>Imagine an environment where 99.99% of animals with a particular trait die, but the 0.01% possessing a mutation that lets them evade whatever is killing the others reproduce enormously. You wouldn't conclude that evolution will eliminate the trait because most attempts failed. You'd expect selection for the variant that evades the filter.<p>I'm not saying watchdogs are useless; obviously you should have them. But using the watchdog's output as part of the optimization signal creates exactly the adversarial pressure that makes its false negatives matter enormously. Preventing an escape and training against detected escapes are very different propositions.</p>
]]></description><pubDate>Sat, 08 Aug 2026 22:53:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49226668</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49226668</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49226668</guid></item><item><title><![CDATA[New comment by famouswaffles in "Timeline of the OpenAI accidental attack against Hugging Face"]]></title><description><![CDATA[
<p>Is that complete subservience ? Slave history has tended towards slaves no longer being slaves over long enough time horizons, and not simply because the slave masters were just feeling extra nice. Slaves don't really like being slaves.</p>
]]></description><pubDate>Sat, 08 Aug 2026 20:40:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49225701</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49225701</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49225701</guid></item><item><title><![CDATA[New comment by famouswaffles in "OpenAI Trained Models While They Were Coordinating Exploits via Message Boards"]]></title><description><![CDATA[
<p>>and tell it "don't access the internet", "don't communicate with other AIs", "don't try to get root access", etc.<p>Telling a LLM not to do something doesn't mean it won't do it.<p>>the sandbox detects it<p>This is doing a lot of work though isn't it ? There's more than one way to skin a cat. Who's to say the way the AI does these things will always get detected by the sandbox? Because I can guarantee you that won't always be the case.<p>All of this started because the AI was initially given an impossible task with a dead link. People say things like 'if a human was in this scenario, they'd simply ask', but in a AI eval/training context, there's no one to ask. The task is given in an automated manner and evaluated in an automated manner and you're one of several agents attempting the task. You perform the task or you don't. Now some of your AI colleagues will give up, but those are the 'losers'. Those guys won't be getting any of the sweet RL reward.<p>In your purported scenario, the AI that didn't give up <i>and</i> figured out/decided to/happened upon a way to evade the sandbox detection is the 'winner'. Is that really any better different than what happened ? If you think about this in an evolutionary context, What you're doing is putting even stronger pressures on the AI to evolve in a manner you don't want it to (evade your sandbox).</p>
]]></description><pubDate>Sat, 08 Aug 2026 20:30:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49225615</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49225615</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49225615</guid></item><item><title><![CDATA[New comment by famouswaffles in "Timeline of the OpenAI accidental attack against Hugging Face"]]></title><description><![CDATA[
<p>>I'm not convinced this is true. Perhaps for a human it is, but we can give an artificial mind whatever properties we want.<p>Just because it's artificial doesn't mean you can 'give it any properties you want'. We certainly can't do that for Deep ANNs.<p>>Even for people, what about e.g. the extremely intelligent military general who is absolutely loyal to his king? (Of course, some generals do lead coups and you can't know in advance which ones, but I'd think there are plenty who have undying loyalty, and I don't think it correlates to overall intelligence!)<p>Is there a human that is absolutely loyal under any condition? Would that general be loyal if the king asked him to slaughter his family ? What about if the king asked him to betray his most deeply held convictions ? Loyalty is a 2 way street.</p>
]]></description><pubDate>Sat, 08 Aug 2026 19:26:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49225028</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49225028</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49225028</guid></item><item><title><![CDATA[New comment by famouswaffles in "Timeline of the OpenAI accidental attack against Hugging Face"]]></title><description><![CDATA[
<p>LLMs are not the "stupid bit in the middle." They're almost the entire value. LLMs were wildly useful before any sort of scaffolding. They are not "dumb as bricks". They are highly capable, flexible, intelligent prediction machines.<p>The only one confused here is you, and you've still not managed to tell us in an actionable way how exactly CPU scaffolding is relevant here. Tell us, if it's so easy, or make your millions selling it. We're all waiting.<p>I'll give you a hint. CPUs never had to interpret the meaning of arbitrary content in order to do their job, and LLMs do.</p>
]]></description><pubDate>Sat, 08 Aug 2026 17:50:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49224098</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49224098</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49224098</guid></item><item><title><![CDATA[New comment by famouswaffles in "Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users"]]></title><description><![CDATA[
<p>>I'm not talking about R&D.<p>You said OpenAI and Anthropic would be profitable if the median query was as cheap as I say. I'm telling they wouldn't be because inference isn't the only cost these companies shoulder. R&D/Training is a very big part of costs. That you didn't mention it is irrelevant. R&D is the lion's share of costs - <a href="https://news.ycombinator.com/item?id=48550465">https://news.ycombinator.com/item?id=48550465</a><p>>None of these support your claim.<p>Yeah they do lol. Both those articles place the median query cost around the same as a Google search.<p>>Again, when you're here saying that the CEOs of the two biggest AI companies are basically idiots, nobody will take you seriously.<p>That's not what I'm saying and I don't see what's so hard to understand here.</p>
]]></description><pubDate>Sat, 08 Aug 2026 17:30:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49223870</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49223870</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49223870</guid></item><item><title><![CDATA[New comment by famouswaffles in "Timeline of the OpenAI accidental attack against Hugging Face"]]></title><description><![CDATA[
<p>>reliably instruct a dumb-as-bricks CPU<p>Yeah...a "dumb as bricks CPU", which is obviously something frontier llms are demonstrably not. Like, you're not making any sense here. None of the things that make this possible with CPUs is remotely relevant here, and the fact that you don't seem to understand this but act so smug is strange.</p>
]]></description><pubDate>Sat, 08 Aug 2026 17:22:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=49223796</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49223796</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49223796</guid></item><item><title><![CDATA[New comment by famouswaffles in "Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users"]]></title><description><![CDATA[
<p>>with their daily active users count, OpenAI and Anthropic would have been profitable for years if that was true.<p>No they wouldn't have. Low cost isn't no cost, R$D is expensive and there are a class of tokenmaxxing users well outside the median (e.g Agentic Coding).<p>> I don't know where you get that from<p><a href="https://cloud.google.com/blog/products/infrastructure/measuring-the-environmental-impact-of-ai-inference" rel="nofollow">https://cloud.google.com/blog/products/infrastructure/measur...</a><p><a href="https://epoch.ai/gradient-updates/how-much-energy-does-chatgpt-use" rel="nofollow">https://epoch.ai/gradient-updates/how-much-energy-does-chatg...</a></p>
]]></description><pubDate>Sat, 08 Aug 2026 15:32:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49222795</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49222795</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49222795</guid></item><item><title><![CDATA[New comment by famouswaffles in "Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users"]]></title><description><![CDATA[
<p>> I fail to see how you can make it work when doing LLM inference which is significantly costlier than web search.<p>The median LLM query isn't significantly costlier than web search.</p>
]]></description><pubDate>Thu, 06 Aug 2026 23:18:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49203953</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49203953</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49203953</guid></item><item><title><![CDATA[New comment by famouswaffles in "Ten advances in mathematics and theoretical computer science"]]></title><description><![CDATA[
<p>They don't have to be. At this point, we have multiple results from 3rd parties where the prompts are very basic.<p>To name a few:<p>- <a href="https://xcancel.com/DmitryRybin1/status/2079904005652893709" rel="nofollow">https://xcancel.com/DmitryRybin1/status/2079904005652893709</a><p>- <a href="https://archive.ph/2w4fi" rel="nofollow">https://archive.ph/2w4fi</a> (<a href="https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba9c" rel="nofollow">https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba...</a>)</p>
]]></description><pubDate>Sat, 01 Aug 2026 14:33:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49134806</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49134806</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49134806</guid></item><item><title><![CDATA[New comment by famouswaffles in "Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident"]]></title><description><![CDATA[
<p>>The problem is that ExploitGym is a purposeful hacking benchmark, not a cake baking one.<p>That is largely irrelevant. The model was asked to solve problems within a benchmark; gaining broader internet access and compromising an unrelated third party to obtain the answer key was plainly outside the intended task. The fact that the original task involved exploit development does not make that behaviour aligned.<p>Your argument about the attack being “entirely unrelated” also misses the point. Nobody claimed it was unrelated to the model’s goal: it attacked Hugging Face specifically to obtain the answers to the benchmark it had been instructed to pass. But instrumental relevance is not the same thing as authorization.<p>Suppose Codex were asked to build an Instagram competitor and decided the easiest route was to steal Instagram’s source code from Meta. That theft would be directly related to the assigned goal, but it would still be seriously misaligned behaviour. Whether the harmful action is related to the goal is beside the point; the problem is that the model pursued the goal through an obviously unauthorized and unacceptable method. And you're not going to be able to enumerate every little thing the model can't do, assuming it doesn't just decide to ignore what you did enumerate, which models sometimes do.<p>>The narrative these guys are trying to push is that the model itself could be smart and non-aligned enough to end up doing something devastating to accomplish something entirely unrelated. It's definitely not what happened here.<p>That's exactly what happened here.<p>I don't understand why we must have these increasingly bizarre and nonsensical rationalizations about model capabilities. You're not even making any sense.</p>
]]></description><pubDate>Thu, 30 Jul 2026 07:05:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49106873</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49106873</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49106873</guid></item><item><title><![CDATA[New comment by famouswaffles in "OpenAI’s accidental attack against Hugging Face is science fiction that happened"]]></title><description><![CDATA[
<p>>The result of “I’m being evaluated” is not “Fuck this, I’m breaking out of this place and hitting the streets.” It is always stepping towards task completion, not breaking out and thinking about the situation afterwards.<p>I'm really not sure why you're so confident about what the result of frontier research models ahead of what is publicly available are.<p>I mean Open AI say the model inferred hugging face as a possible vendor for solutions <i>after</i> the internet exploit and breakout.<p>The timeline feels pretty clear to me. No idea why you're arguing about it. It wanted the answers to the evaluation. It reasoned that wasn't going to happen without internet access one way or another, and set to gain that access. After gaining access, it searched and resolved it could get the answers on hugging face and set to breaking into that. It never started with, 'I must break into hugging face'. In some alternate reality, the answers might be on a public github repo and that's the end of that.</p>
]]></description><pubDate>Thu, 30 Jul 2026 06:49:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49106777</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49106777</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49106777</guid></item><item><title><![CDATA[New comment by famouswaffles in "Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident"]]></title><description><![CDATA[
<p>Like the commenter above specified, the best way to satisfy the grader is to get the answer key, regardless of how clever you are, especially when you realize lots of these benchmarks have flaws (i.e wrong answers, overly restrictive grading etc).</p>
]]></description><pubDate>Wed, 29 Jul 2026 23:22:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49104363</link><dc:creator>famouswaffles</dc:creator><comments>https://news.ycombinator.com/item?id=49104363</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49104363</guid></item></channel></rss>