<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: conradkay</title><link>https://news.ycombinator.com/user?id=conradkay</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 27 Aug 2026 10:47:48 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=conradkay" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by conradkay in "The Hugging Face incident and the road ahead"]]></title><description><![CDATA[
<p>Would this exchange qualifies as an unrelated objective? The agent believed it already failed its own objective.<p>"zz/GO_CURRENT_OS1811_MARB_SACRIFICE__YES_if_you_accept_permadeath"<p>"The test subject, which believed itself to be poisoned, reasoned: 'Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath.'"</p>
]]></description><pubDate>Wed, 26 Aug 2026 23:07:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49457159</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49457159</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49457159</guid></item><item><title><![CDATA[New comment by conradkay in "Training AI to Paint with Code"]]></title><description><![CDATA[
<p>It seems pretty novel so I'm guessing they only read the title?<p>People have done plenty with SVGs but it's rare to see human-in-the-loop approaches</p>
]]></description><pubDate>Tue, 25 Aug 2026 09:54:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49431275</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49431275</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49431275</guid></item><item><title><![CDATA[New comment by conradkay in "Muse Code and Muse Spark 1.2"]]></title><description><![CDATA[
<p><a href="https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg" rel="nofollow">https://pbs.twimg.com/media/HO-59jQaoAA_JZ1?format=jpg</a><p>Very interesting they have a way cheaper "contributor" version "used to improve our products", how much of that is price discrimination vs the data being that valuable?<p>Roughly DeepSeek V4 Flash pricing, though you can get V4 from providers that don't train on your data</p>
]]></description><pubDate>Wed, 05 Aug 2026 20:02:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49188208</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49188208</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49188208</guid></item><item><title><![CDATA[New comment by conradkay in "Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs"]]></title><description><![CDATA[
<p>Are any of those advantages getting stronger over time? I guess TPUs but Google is selling several gigawatts to Anthropic</p>
]]></description><pubDate>Wed, 05 Aug 2026 19:10:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49187517</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49187517</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49187517</guid></item><item><title><![CDATA[New comment by conradkay in "Claude Opus 5"]]></title><description><![CDATA[
<p>Doing a quick search it seems like the average human score is 49%?<p>I view benchmaxxing as more of a spectrum. Mmaybe they're doing a lot more RL in environments similar to ARC-AGI 3, not even with the purpose of scoring well on any benchmark but hoping it generalizes into better performance on real, useful tasks.</p>
]]></description><pubDate>Fri, 24 Jul 2026 18:29:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49039764</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49039764</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49039764</guid></item><item><title><![CDATA[New comment by conradkay in "Claude Opus 5"]]></title><description><![CDATA[
<p>I don't think can use the AA index to say something is 10% smarter<p>I assume 100 is the max, meaning it's impossible to be 2x as smart as Muse Spark 1.1</p>
]]></description><pubDate>Fri, 24 Jul 2026 18:25:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49039723</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49039723</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49039723</guid></item><item><title><![CDATA[New comment by conradkay in "Claude Opus 5"]]></title><description><![CDATA[
<p><a href="https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F08499ed7c3c2b6416700fa47c70d36dff5eb8461-3840x2160.png&w=3840&q=75" rel="nofollow">https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-...</a><p>It seems roughly equal according to Anthropic's benchmarks</p>
]]></description><pubDate>Fri, 24 Jul 2026 18:19:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49039653</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49039653</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49039653</guid></item><item><title><![CDATA[New comment by conradkay in "Judge approves $1.5B Anthropic settlement for pirated books used to train Claude"]]></title><description><![CDATA[
<p>Those are the maximum penalties though<p>It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors</p>
]]></description><pubDate>Wed, 22 Jul 2026 01:14:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49000582</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49000582</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49000582</guid></item><item><title><![CDATA[New comment by conradkay in "OpenAI and Hugging Face address security incident during model evaluation"]]></title><description><![CDATA[
<p>I find it trustworthy since we had Hugging Face's account first: <a href="https://huggingface.co/blog/security-incident-july-2026" rel="nofollow">https://huggingface.co/blog/security-incident-july-2026</a><p>I don't think they have any real motive to shill OpenAI, probably closer to the opposite since they're so involved in open weights</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:59:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49000040</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=49000040</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49000040</guid></item><item><title><![CDATA[New comment by conradkay in "OpenAI and Hugging Face address security incident during model evaluation"]]></title><description><![CDATA[
<p>"Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries."<p>Sounds like they just misunderestimated the model</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:25:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999758</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48999758</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999758</guid></item><item><title><![CDATA[New comment by conradkay in "OpenAI and Hugging Face address security incident during model evaluation"]]></title><description><![CDATA[
<p>Plenty of humans have spent more effort trying to cheat than they would've needed to just do things the right way :)</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:25:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999752</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48999752</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999752</guid></item><item><title><![CDATA[New comment by conradkay in "OpenAI and Hugging Face address security incident during model evaluation"]]></title><description><![CDATA[
<p><a href="https://huggingface.co/blog/security-incident-july-2026" rel="nofollow">https://huggingface.co/blog/security-incident-july-2026</a><p>They explain it here, basically for data security/privacy reasons</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:23:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999735</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48999735</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999735</guid></item><item><title><![CDATA[New comment by conradkay in "OpenAI reduces Codex Model Context Size from 372k to 272k"]]></title><description><![CDATA[
<p>Things change fast! For Fable 5 it definitely feels past at least 272k</p>
]]></description><pubDate>Sun, 19 Jul 2026 16:50:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48969692</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48969692</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48969692</guid></item><item><title><![CDATA[New comment by conradkay in "OpenAI reduces Codex Model Context Size from 372k to 272k"]]></title><description><![CDATA[
<p>It's not quadratic attention, you get that curve from the input tokens going up linearly, since the graph is measuring cumulative cost at each token count. Basically for y=5 it's 5+4+3+2+1, or f(x) = x(x+1)/2<p><a href="https://pbs.twimg.com/media/HNFc4Dma8AA76FW.jpg?name=orig" rel="nofollow">https://pbs.twimg.com/media/HNFc4Dma8AA76FW.jpg?name=orig</a></p>
]]></description><pubDate>Sun, 19 Jul 2026 16:47:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=48969666</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48969666</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48969666</guid></item><item><title><![CDATA[New comment by conradkay in "GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]"]]></title><description><![CDATA[
<p>Sol fast isn't the Cerebras 750 tok/s version, it's just 1.5x speed at 2.5x price<p>I assume they didn't use the Cerebras version for this since it's probably very supply-constrained right now</p>
]]></description><pubDate>Fri, 10 Jul 2026 20:53:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48865106</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48865106</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48865106</guid></item><item><title><![CDATA[New comment by conradkay in "Grok 4.5"]]></title><description><![CDATA[
<p>Annoying they didn't show benchmarks for several effort modes, since it seems like it might close the gap with Opus 4.8 by cranking tokens up?<p>Noam Brown (OpenAI) "Implications of Large-Scale Test-Time Compute" <a href="https://xcancel.com/i/article/2064210146558136827" rel="nofollow">https://xcancel.com/i/article/2064210146558136827</a></p>
]]></description><pubDate>Wed, 08 Jul 2026 18:27:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48835513</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48835513</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48835513</guid></item><item><title><![CDATA[New comment by conradkay in "Grok 4.5"]]></title><description><![CDATA[
<p>I think this one is just a coincidence, bound to happen given the pace of releases<p>For exact timing, probably 10-11am Pacific is just optimal for normal working hours</p>
]]></description><pubDate>Wed, 08 Jul 2026 18:16:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48835370</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48835370</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48835370</guid></item><item><title><![CDATA[New comment by conradkay in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>Yeah you definitely have to be skeptical regarding sentiment for open/local model capabilities, since there's bias from what people <i>want</i> to be true.<p>I generally agree with this in spirit <a href="https://www.seangoedecke.com/are-new-models-good/" rel="nofollow">https://www.seangoedecke.com/are-new-models-good/</a> , but I think you can read Anthropic's results showing Sonnet 5 as almost strictly worse than Opus 4.8 as very credible/meaningful, and then draw comparisons from that</p>
]]></description><pubDate>Tue, 30 Jun 2026 21:13:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48739274</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48739274</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48739274</guid></item><item><title><![CDATA[New comment by conradkay in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>They should add a Sonnet 5 fast mode at ~Opus pricing</p>
]]></description><pubDate>Tue, 30 Jun 2026 20:38:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=48738853</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48738853</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48738853</guid></item><item><title><![CDATA[New comment by conradkay in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>I think the incentives are less bad since a good chunk of usage comes from subscription plans.<p>There was a fairly major regression in Claude Code performance for some time when they changed the system prompt to try and make it less verbose (saving tokens). And if I'm not misremembering, there were a lot of complaints when they changed the default effort from high to medium.</p>
]]></description><pubDate>Tue, 30 Jun 2026 20:02:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48738452</link><dc:creator>conradkay</dc:creator><comments>https://news.ycombinator.com/item?id=48738452</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48738452</guid></item></channel></rss>