<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: samusiam</title><link>https://news.ycombinator.com/user?id=samusiam</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 24 Sep 2026 18:51:24 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=samusiam" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by samusiam in "Contrastive Language Models"]]></title><description><![CDATA[
<p>I think System 1 is a great term. System 1 is fast, intuitive, and automatic. It describes decision-making that happens without explicit reasoning or deliberation. System 2 is the opposite: it's the more deliberate, "executive functioning" side of cognition -- the part that reasons through a problem before arriving at an answer. That's also what state-of-the-art LLMs do before they respond. Jev doesn’t do that kind of reasoning. It just decides.</p>
]]></description><pubDate>Thu, 24 Sep 2026 11:16:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49829036</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49829036</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49829036</guid></item><item><title><![CDATA[New comment by samusiam in "Strands Harness"]]></title><description><![CDATA[
<p>> But the moment you build your own agent, you’re on your own. It’s tricky wiring up the right primitives just well enough to match that “it just worked” feeling.<p>It's wild to me to claim that it's tricky to customize one of these harnesses and for that to be the entire justification for an entirely different harness.<p>It's really not that hard. If you want to reduce costs then all you need to do is practice delegation: instead of using the strong model, all the time to do everything, instead, you have the stronger model delegate well-defined tasks to a weaker model. Patterns like these are really easy to wire up.</p>
]]></description><pubDate>Wed, 23 Sep 2026 15:37:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49817745</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49817745</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49817745</guid></item><item><title><![CDATA[New comment by samusiam in "Jev in 25 Lines of Python"]]></title><description><![CDATA[
<p>It wasn't</p>
]]></description><pubDate>Wed, 23 Sep 2026 10:54:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49814157</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49814157</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49814157</guid></item><item><title><![CDATA[New comment by samusiam in "I asked Meta’s Muse for its filesystem and it sent me 6.8GB"]]></title><description><![CDATA[
<p>You call it a guess, I call it a hypothesis.</p>
]]></description><pubDate>Wed, 23 Sep 2026 01:37:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49810601</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49810601</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49810601</guid></item><item><title><![CDATA[New comment by samusiam in "The UV index is not the warm sensation of sunlight on bare skin"]]></title><description><![CDATA[
<p>After you sure that isn't because your skin is moist where sunblock was recently applied, and therefore cooler?</p>
]]></description><pubDate>Wed, 23 Sep 2026 01:34:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49810579</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49810579</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49810579</guid></item><item><title><![CDATA[New comment by samusiam in "GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?"]]></title><description><![CDATA[
<p>To me, one interesting piece of analysis was whether Luna had benefit on top of Astra -- i.e., running both and synthesizing their findings. But even with that, it raises the question whether running a second Astra pass, or Sol, or even a model from another family (GLM? Fable?) would deliver even more additive benefit.</p>
]]></description><pubDate>Tue, 15 Sep 2026 00:03:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49705954</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49705954</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49705954</guid></item><item><title><![CDATA[New comment by samusiam in "GPT-5.6 Luna vs. GPT-6 Astra: Is a $1.20 Model Good Enough for Code Review?"]]></title><description><![CDATA[
<p>My thoughts exactly. That 22pp gap is massive, and it seems to me Astra is worth every extra penny.</p>
]]></description><pubDate>Mon, 14 Sep 2026 23:58:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49705923</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49705923</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49705923</guid></item><item><title><![CDATA[New comment by samusiam in "Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1"]]></title><description><![CDATA[
<p>Which is pretty much a useless (i.e., saturated, contaminated) benchmark now.</p>
]]></description><pubDate>Thu, 10 Sep 2026 22:55:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49651244</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=49651244</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49651244</guid></item><item><title><![CDATA[New comment by samusiam in "Reddit Starts Blocking Mobile Website, Pushing Users to App Instead"]]></title><description><![CDATA[
<p>I'm still using Sync, going on three years since they banned third party apps, and third party apps are still the best way to experience the site.</p>
]]></description><pubDate>Tue, 12 May 2026 11:28:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48106737</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=48106737</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48106737</guid></item><item><title><![CDATA[New comment by samusiam in "OpenWarp"]]></title><description><![CDATA[
<p>I don't see anything asking for a credit card.</p>
]]></description><pubDate>Fri, 01 May 2026 12:06:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=47973816</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47973816</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47973816</guid></item><item><title><![CDATA[New comment by samusiam in "Over-editing refers to a model modifying code beyond what is necessary"]]></title><description><![CDATA[
<p>In my experience this "gaming" behavior is easily caught by just asking another agent (could just be another session of Claude Code) to review the code changes.</p>
]]></description><pubDate>Fri, 24 Apr 2026 10:59:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=47888476</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47888476</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47888476</guid></item><item><title><![CDATA[New comment by samusiam in "An update on recent Claude Code quality reports"]]></title><description><![CDATA[
<p>For idle sessions I would MUCH rather pay the cost in tokens than reduced quality. Frankly, it's shocking to me that you would make that trade-off for users without their knowledge or consent.</p>
]]></description><pubDate>Fri, 24 Apr 2026 10:41:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=47888323</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47888323</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47888323</guid></item><item><title><![CDATA[New comment by samusiam in "The threat is comfortable drift toward not understanding what you're doing"]]></title><description><![CDATA[
<p>That's true, but the "AI bubble bursts" scenario is usually tied to Western investors getting essentially margin-called. If that happens, the CCP won't suddenly stop their investment; Chinese models will most likely continue developing.</p>
]]></description><pubDate>Sun, 05 Apr 2026 13:51:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=47649455</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47649455</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47649455</guid></item><item><title><![CDATA[New comment by samusiam in "Caveman: Why use many token when few token do trick"]]></title><description><![CDATA[
<p>> I have no clue if this claim holds, but alas, just pretending they did not address the obvious criticism, while they did, is at the very least pretty lazy.<p>But they didn't address the criticism. "cutting ~75% of tokens while keeping full technical accuracy" is an empirical claim for which no evidence was provided.</p>
]]></description><pubDate>Sun, 05 Apr 2026 13:16:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=47649137</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47649137</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47649137</guid></item><item><title><![CDATA[New comment by samusiam in "The threat is comfortable drift toward not understanding what you're doing"]]></title><description><![CDATA[
<p>> Aren't they currently propped up by investor money?<p>Are Chinese model shops propped up by investor money? Is Google?<p>Open weights models are only 6 months behind SOTA. If new model development suddenly stopped, and today's SOTA models suddenly disappeared, we would still have access to capable agents.</p>
]]></description><pubDate>Sun, 05 Apr 2026 13:04:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=47649007</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47649007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47649007</guid></item><item><title><![CDATA[New comment by samusiam in "Qwen3.6-Plus: Towards Real World Agents"]]></title><description><![CDATA[
<p>These OSS model makers need to stop benchmarking against old models. Showing how it performs against Opus 4.5, GLM-5 when we have Opus 4.6 and GLM-5.1 just tells me that it's not comparable to SOTA.</p>
]]></description><pubDate>Thu, 02 Apr 2026 12:03:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=47613295</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47613295</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47613295</guid></item><item><title><![CDATA[New comment by samusiam in "The revenge of the data scientist"]]></title><description><![CDATA[
<p>I think there's a lot of methodological expertise that goes into collecting good eval data. For example, in many cases you need human labelers with the right expertise, well designed tasks, well defined constructs, and you need to hit interrater agreement targets and troubleshoot when you don't. Good label data is a prerequisite to the stuff that can probably be automated by the AI agent (improving the system to optimize a metric measured against ground truth labels). Data scientists and research scientists are more likely to have this skillset. And it takes time to pick up and learn the nuances.</p>
]]></description><pubDate>Wed, 01 Apr 2026 23:48:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=47608174</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47608174</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47608174</guid></item><item><title><![CDATA[New comment by samusiam in "Do your own writing"]]></title><description><![CDATA[
<p>Well you're free to disagree but my experience has been counter to your position. I write both code and research / technical documentation. The quality of what the LLM produces is limited by the quality of ideas I give it initially (mind you, this is just a starting point), and the quality of my review of its output.</p>
]]></description><pubDate>Wed, 01 Apr 2026 16:51:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=47603396</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47603396</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47603396</guid></item><item><title><![CDATA[New comment by samusiam in "Claude Code Unpacked : A visual guide"]]></title><description><![CDATA[
<p>You're complaining about vibe coding while also complaining about how you "feel" about the code. Do you see the irony in that?</p>
]]></description><pubDate>Wed, 01 Apr 2026 13:55:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=47600937</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47600937</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47600937</guid></item><item><title><![CDATA[New comment by samusiam in "Claude Code Unpacked : A visual guide"]]></title><description><![CDATA[
<p>I haven't seen the scrolling glitch in months, where previously it was happening multiple times a day. Also haven't seen anyone complain about it in quite some time. Pretty sure they have resolved that.</p>
]]></description><pubDate>Wed, 01 Apr 2026 13:53:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=47600903</link><dc:creator>samusiam</dc:creator><comments>https://news.ycombinator.com/item?id=47600903</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47600903</guid></item></channel></rss>