<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: rbuccigrossi</title><link>https://news.ycombinator.com/user?id=rbuccigrossi</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 29 Jul 2026 20:28:42 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=rbuccigrossi" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by rbuccigrossi in "ARC-AGI Leaderboard"]]></title><description><![CDATA[
<p>No, I believe you are misunderstanding the quote. Each “question” in ARC-AGI-3 is a game that has hidden rules that you can understand if you look at the game board long enough. This quote means that Opus 5 is looking at the game board, figuring out the rules, and writing out the rules before it makes a single move. You can do the same thing if you go to the ARC-AGI-3 website and try some of the games.</p>
]]></description><pubDate>Sun, 26 Jul 2026 00:12:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49053210</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=49053210</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49053210</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Is AI Progress Real? Four Independent Metrics Show It"]]></title><description><![CDATA[
<p>The others are an IQ test (TrackingAI), a test of Ph.D. questions across multiple domains (Humanity's Last Exam), and graphical pattern matching (ARC-AGI-2).<p>What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the <i>behavior</i> of the resulting improvement curves are extremely close.<p>That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.</p>
]]></description><pubDate>Mon, 20 Jul 2026 03:20:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48974007</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48974007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48974007</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Is AI Progress Real? Four Independent Metrics Show It"]]></title><description><![CDATA[
<p>The text is my transcript of the video formatted by Claude and with sources added at the end.<p>I do use deep research (across Gemini, ChatGPT, and Claude) to gather background and ideas, and Claude for editing. I started machine learning and computer vision research back in 1995, studied NLP, and dove into LLMs with GPT-2, so I wouldn't be surprised if my writing has been deeply influenced by AI.</p>
]]></description><pubDate>Mon, 20 Jul 2026 03:17:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=48973993</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48973993</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48973993</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Is AI Progress Real? Four Independent Metrics Show It"]]></title><description><![CDATA[
<p>Author here. In short there are 4 different metrics:<p>- METR's time horizon<p>- TrackingAI's offline cognitive test<p>- Humanity's Last Exam, and<p>- ARC-AGI-2<p>that have lasted longer than 2 years (though ARC-AGI-2 is now saturated).<p>When plotted in the linear domain, they all have an exponential (hockey-shaped) curve, but the interesting thing is that the bend happens right at Q4 of 2025 (right when Gemini 3, Opus 4.5, GPT 5.2 all come out).</p>
]]></description><pubDate>Mon, 20 Jul 2026 03:13:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48973973</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48973973</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48973973</guid></item><item><title><![CDATA[Is AI Progress Real? Four Independent Metrics Show It]]></title><description><![CDATA[
<p>Article URL: <a href="https://skepticcto.substack.com/p/is-ai-progress-real-a-skepticcto">https://skepticcto.substack.com/p/is-ai-progress-real-a-skepticcto</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48971910">https://news.ycombinator.com/item?id=48971910</a></p>
<p>Points: 6</p>
<p># Comments: 6</p>
]]></description><pubDate>Sun, 19 Jul 2026 21:39:09 +0000</pubDate><link>https://skepticcto.substack.com/p/is-ai-progress-real-a-skepticcto</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48971910</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48971910</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Local AI Hardware: Break Even in 2.6 Years?"]]></title><description><![CDATA[
<p>In short, running a $3,299 GMKtek EVO-X2 (Ryzen AI Max+395 with 198 GB) 24/7 with the Gemma 4 26B-A4B model, being as generous as possible, only saves you $1,279.07/year in inference costs. (120 t/s for output tokens at $0.34/M.)<p>So, that's how you get a break even at 2.6 years...</p>
]]></description><pubDate>Fri, 29 May 2026 21:21:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48329447</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48329447</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48329447</guid></item><item><title><![CDATA[Local AI Hardware: Break Even in 2.6 Years?]]></title><description><![CDATA[
<p>Article URL: <a href="https://skepticcto.com/news/updates/2026/05/29/LocalAI.html">https://skepticcto.com/news/updates/2026/05/29/LocalAI.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48329356">https://news.ycombinator.com/item?id=48329356</a></p>
<p>Points: 3</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 29 May 2026 21:14:11 +0000</pubDate><link>https://skepticcto.com/news/updates/2026/05/29/LocalAI.html</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48329356</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48329356</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Show HN: Decoding the Language Machine – AI video series and CC repo"]]></title><description><![CDATA[
<p>You're first :D</p>
]]></description><pubDate>Tue, 26 May 2026 21:34:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48286291</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48286291</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48286291</guid></item><item><title><![CDATA[Show HN: Decoding the Language Machine – AI video series and CC repo]]></title><description><![CDATA[
<p>Hi HN! I released 3 parts of an educational video series (out of 6 planned), paired with a GitHub repository containing scripts and artifacts (released under Creative Commons).<p>- Main Site: <a href="https://skepticcto.com/" rel="nofollow">https://skepticcto.com/</a> (includes related AI news articles)<p>- Code/Artifacts: <a href="https://github.com/SkepticCTO/decoding_the_language_machine" rel="nofollow">https://github.com/SkepticCTO/decoding_the_language_machine</a><p>- YouTube Channel: <a href="https://www.youtube.com/@SkepticCTO" rel="nofollow">https://www.youtube.com/@SkepticCTO</a><p>I’m a 21-year CTO, Ph.D. in CS (U Penn, 1999 in computer vision and ML), and a PI in the NIST AI Safety Initiative Consortium. I spent a 4-month sabbatical making this because I wanted to demystify how LLMs work through a historical perspective (starting in 1948 with Claude Shannon) and scientific skepticism.<p>The project is old enough to be fleshed out, but young enough to be able to pivot. Is it useful? What would you like to see? I look forward to questions and feedback.</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48279163">https://news.ycombinator.com/item?id=48279163</a></p>
<p>Points: 2</p>
<p># Comments: 2</p>
]]></description><pubDate>Tue, 26 May 2026 12:56:16 +0000</pubDate><link>https://github.com/SkepticCTO/decoding_the_language_machine</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=48279163</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48279163</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Threatening AI Does Not Make It More Useful. Why Sergey Brin Is Wrong"]]></title><description><![CDATA[
<p>We work in the arena of automated AI workflows where consistency of success is vital. When you threaten an LLM you are drawing the LLM into the texts where threats occur (flame wars, parody, etc.). So intuitively you would expect it to work sometimes, but also fail with even more ardent refusal (increasing the variance of success).<p>Jailbreak approaches like "Bad Likert Judge" ( <a href="https://unit42.paloaltonetworks.com/multi-turn-technique-jailbreaks-llms/" rel="nofollow">https://unit42.paloaltonetworks.com/multi-turn-technique-jai...</a> ) and similar persuasive techniques (see <a href="https://xthemadgenius.medium.com/how-persuasion-techniques-can-jailbreak-language-models-ec9208b9ca49" rel="nofollow">https://xthemadgenius.medium.com/how-persuasion-techniques-c...</a> ) move the text domain to more policy, analysis, or scientific papers, where deeper analysis, discussion, and compliance is the norm.<p>So I'm curious about the extremes (variance) of success with threatening vs. polite discussion, but I haven't seen direct research on that.</p>
]]></description><pubDate>Tue, 17 Jun 2025 19:51:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=44303181</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=44303181</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44303181</guid></item><item><title><![CDATA[New comment by rbuccigrossi in "Threatening AI Does Not Make It More Useful. Why Sergey Brin Is Wrong"]]></title><description><![CDATA[
<p>Treating an LLM with respect is not about pretending it has feelings; it’s about understanding that every word in your prompt is a signal that shifts the probabilistic landscape from which the model draws its answer. It’s about probability, not personality.</p>
]]></description><pubDate>Tue, 17 Jun 2025 19:34:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=44302979</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=44302979</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44302979</guid></item><item><title><![CDATA[Threatening AI Does Not Make It More Useful. Why Sergey Brin Is Wrong]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.tcg.com/blog/does-being-rude-to-ai-make-it-more-useful-why-sergey-brin-is-wrong/">https://www.tcg.com/blog/does-being-rude-to-ai-make-it-more-useful-why-sergey-brin-is-wrong/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=44302978">https://news.ycombinator.com/item?id=44302978</a></p>
<p>Points: 5</p>
<p># Comments: 3</p>
]]></description><pubDate>Tue, 17 Jun 2025 19:34:38 +0000</pubDate><link>https://www.tcg.com/blog/does-being-rude-to-ai-make-it-more-useful-why-sergey-brin-is-wrong/</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=44302978</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44302978</guid></item><item><title><![CDATA[“End of Support” Does Not Mean “End of Life” for Open Source Projects]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.tcg.com/blog/end-of-support-does-not-mean-end-of-life-for-open-source-projects/">https://www.tcg.com/blog/end-of-support-does-not-mean-end-of-life-for-open-source-projects/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=23853479">https://news.ycombinator.com/item?id=23853479</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 15 Jul 2020 23:05:05 +0000</pubDate><link>https://www.tcg.com/blog/end-of-support-does-not-mean-end-of-life-for-open-source-projects/</link><dc:creator>rbuccigrossi</dc:creator><comments>https://news.ycombinator.com/item?id=23853479</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=23853479</guid></item></channel></rss>