<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: kakugawa</title><link>https://news.ycombinator.com/user?id=kakugawa</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 26 Sep 2026 03:07:44 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=kakugawa" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by kakugawa in "Ollaya – Ollama for open-source, Jev-style decision models"]]></title><description><![CDATA[
<p>Jev's value becomes more apparent when the task is a moving target. eg an auto-mode classifier.</p>
]]></description><pubDate>Fri, 25 Sep 2026 21:53:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49850457</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49850457</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49850457</guid></item><item><title><![CDATA[New comment by kakugawa in "Opus 5.5 is good at explainer videos"]]></title><description><![CDATA[
<p>Here are a couple examples from r/ClaudeAI:<p>Made entirely with Opus 5.5 + $3.21 of OpenRouter API usageClaude Code Workflow<p><a href="https://www.reddit.com/r/ClaudeAI/comments/1wogab3/" rel="nofollow">https://www.reddit.com/r/ClaudeAI/comments/1wogab3/</a><p>Jaw literally dropped. I ran the prompt from the "Made entirely with Opus 5.5" post on my own project. Here's what Claude Code made on its own for about $4.Claude Code Workflow<p><a href="https://www.reddit.com/r/ClaudeAI/comments/1wovwao/" rel="nofollow">https://www.reddit.com/r/ClaudeAI/comments/1wovwao/</a></p>
]]></description><pubDate>Thu, 24 Sep 2026 21:38:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49837061</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49837061</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49837061</guid></item><item><title><![CDATA[New comment by kakugawa in "AI recursive self-improvement might not come so quickly after all"]]></title><description><![CDATA[
<p>They can only do it in the (narrow) domains that are verifiable.</p>
]]></description><pubDate>Sun, 13 Sep 2026 22:35:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49689434</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49689434</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49689434</guid></item><item><title><![CDATA[New comment by kakugawa in "OpenAI Agents API"]]></title><description><![CDATA[
<p>I assume you'd develop via the SDK, then deploy it via the API.</p>
]]></description><pubDate>Thu, 10 Sep 2026 21:33:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49650446</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49650446</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49650446</guid></item><item><title><![CDATA[New comment by kakugawa in "GPT-6 Astra on robot arms"]]></title><description><![CDATA[
<p>It's not a coincidence that Waymo started becoming viable after GPT-3.</p>
]]></description><pubDate>Sun, 06 Sep 2026 03:11:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49582964</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49582964</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49582964</guid></item><item><title><![CDATA[New comment by kakugawa in "Show HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x"]]></title><description><![CDATA[
<p>With all the hype around the latest Gemini 3.8 Flash/Cyber release, will Antigravity CLI [1] be supported?<p>1/ <a href="https://antigravity.google/product/antigravity-cli" rel="nofollow">https://antigravity.google/product/antigravity-cli</a></p>
]]></description><pubDate>Wed, 02 Sep 2026 18:04:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49540040</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49540040</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49540040</guid></item><item><title><![CDATA[New comment by kakugawa in "Solaris, Our Interface World Models"]]></title><description><![CDATA[
<p><a href="https://www.youtube.com/watch?v=c2yCePPnrSA" rel="nofollow">https://www.youtube.com/watch?v=c2yCePPnrSA</a><p>Their analogy is that a VLM responds to text, like their interface model responds to clicks. There is no UI (just images), and based on your clicks the model infers your intent and adapts the "UI" in response. So, instead of inferring intent from an information-dense input (text), they do it w/ just mouse-based gestures? I would love to see how this holds up in practice.<p>A fun little anecdote @ 54s in the video: "It becomes whatever you ask of it. And no two interactions are ever the same."</p>
]]></description><pubDate>Mon, 31 Aug 2026 20:17:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49514390</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49514390</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49514390</guid></item><item><title><![CDATA[New comment by kakugawa in "AutoSaddler: Automatic Harness Optimization"]]></title><description><![CDATA[
<p><a href="https://autosaddler-projectpage.github.io/" rel="nofollow">https://autosaddler-projectpage.github.io/</a><p><a href="https://github.com/microsoft/AutoSaddler" rel="nofollow">https://github.com/microsoft/AutoSaddler</a></p>
]]></description><pubDate>Fri, 28 Aug 2026 18:08:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49482315</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49482315</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49482315</guid></item><item><title><![CDATA[New comment by kakugawa in "Show HN: We built open OpenRouter that turns usage into a better model"]]></title><description><![CDATA[
<p>They make money on enterprise plans: <a href="https://www.experientiallabs.ai/pricing#enterprise">https://www.experientiallabs.ai/pricing#enterprise</a><p>Look at the Intelligence features in the Enterprise plan:<p>* Per-prompt model optimization<p>* Caching<p>* A model you own, trained on your traffic</p>
]]></description><pubDate>Fri, 28 Aug 2026 03:58:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49474187</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49474187</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49474187</guid></item><item><title><![CDATA[New comment by kakugawa in "Small Models Have Arrived"]]></title><description><![CDATA[
<p>FrontierCode is prob the closest. [1] It's closed source (so no direct benchmaxxing), and it was calibrated by 20+ open source maintainers. It shows Opus 5 (medium), beating out the other reasoning levels by a large margin. i.e. Opus 5 w/ higher reasoning levels actually reduces performance. [2]<p>However, you'll have to gauge for yourself how closely their tasks resemble your tasks.<p>1/ <a href="https://cognition.com/blog/frontier-code" rel="nofollow">https://cognition.com/blog/frontier-code</a><p>2/ <a href="https://cognition.com/frontiercode" rel="nofollow">https://cognition.com/frontiercode</a></p>
]]></description><pubDate>Fri, 28 Aug 2026 00:06:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49472821</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49472821</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49472821</guid></item><item><title><![CDATA[New comment by kakugawa in "I Analyzed 163K Lines of Kuzu's Codebase. Here's Why Apple Wanted It"]]></title><description><![CDATA[
<p><a href="https://archive.is/iV3sk" rel="nofollow">https://archive.is/iV3sk</a></p>
]]></description><pubDate>Thu, 20 Aug 2026 21:38:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49380630</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49380630</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49380630</guid></item><item><title><![CDATA[New comment by kakugawa in "Cursor launches Origin, GitHub alternative"]]></title><description><![CDATA[
<p>I believe Entire.io is also trying to build a Github replacement. (Not affiliated, but I use the Entire.io CLI and I find it useful.)</p>
]]></description><pubDate>Mon, 17 Aug 2026 21:48:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49338150</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49338150</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49338150</guid></item><item><title><![CDATA[New comment by kakugawa in "Anthropic: Introducing The Conceptual Reasoning Index"]]></title><description><![CDATA[
<p>"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.</p>
]]></description><pubDate>Thu, 13 Aug 2026 14:56:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49286999</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49286999</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49286999</guid></item><item><title><![CDATA[New comment by kakugawa in "The End of No Code"]]></title><description><![CDATA[
<p>I feel like there's going to be a migration to a new interface paradigm that's balanced between agentic use and human review (with light editing). The current low/no code UIs are human-centric (visual, point-and-click). Going full agent-centric greatly increases the chance of slop code. You want an interface that an agent can natively drive, but that a human can still reasonably review and lightly edit. Probably a domain-specific text-based interface.</p>
]]></description><pubDate>Thu, 06 Aug 2026 18:51:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49200731</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49200731</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49200731</guid></item><item><title><![CDATA[New comment by kakugawa in "Superlogical"]]></title><description><![CDATA[
<p>They can at least deliver baseline value, because it's all text-based. Do agree that it'll become more complex, when/if they decide to create custom protocols.</p>
]]></description><pubDate>Thu, 30 Jul 2026 00:08:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49104663</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49104663</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49104663</guid></item><item><title><![CDATA[New comment by kakugawa in "Show HN: Claude-thermos keeps your Claude session warm for you"]]></title><description><![CDATA[
<p>The prefix cache is a resource shared by all users. This is basically a tragedy of the commons.</p>
]]></description><pubDate>Thu, 23 Jul 2026 23:48:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=49029616</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=49029616</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49029616</guid></item><item><title><![CDATA[New comment by kakugawa in "Separating signal from noise in coding evaluations"]]></title><description><![CDATA[
<p>The more subtle point is that there's a gap between the task and its verification. e.g. if you have an open-ended / under-specified prompt, the verification needs to be able to handle all potential solutions.<p>So you can have a very narrow task prompt that's easy to verify (but likely too simple of a challenge). Or a more realistic task prompt that's much harder to verify. And likely harder to both build the robust verifier and run it cheaply.</p>
]]></description><pubDate>Wed, 08 Jul 2026 22:17:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48838130</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=48838130</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48838130</guid></item><item><title><![CDATA[New comment by kakugawa in "Computer use in Gemini 3.5 Flash"]]></title><description><![CDATA[
<p>Antigravity CLI (which replaced Gemini CLI):<p><a href="https://antigravity.google/product/antigravity-cli" rel="nofollow">https://antigravity.google/product/antigravity-cli</a></p>
]]></description><pubDate>Wed, 24 Jun 2026 23:29:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48666807</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=48666807</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48666807</guid></item><item><title><![CDATA[New comment by kakugawa in "Qwen-AgentWorld: Language World Models for General Agents"]]></title><description><![CDATA[
<p>Here's the demo: <a href="https://docs.qwenlm.ai/resources/mlu56_demo.html" rel="nofollow">https://docs.qwenlm.ai/resources/mlu56_demo.html</a><p>Here's the description of the world model prompt for the web domain: "A precise GUI state simulator — given the current screen (as HTML) and a user action, predicts the exact next screen as a complete, self-contained HTML document." (You can click the world model prompt box to expand it and see the full prompt.)<p>So the world model generates the current state (an html document), an agent tells it what action it wants to perform, the world model generates the next state (another html document).<p>The other domains are similar, but w/ domain-specific nuance.</p>
]]></description><pubDate>Wed, 24 Jun 2026 07:57:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48656638</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=48656638</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48656638</guid></item><item><title><![CDATA[New comment by kakugawa in "The Coming Loop"]]></title><description><![CDATA[
<p>How much does /goal actually help? In auto mode, I've tried using and not using /goal and I haven't felt a difference.<p><a href="https://code.claude.com/docs/en/goal#how-evaluation-works" rel="nofollow">https://code.claude.com/docs/en/goal#how-evaluation-works</a><p>> /goal is a wrapper around a session-scoped prompt-based Stop hook. Each time Claude finishes a turn, the condition and the conversation so far are sent to your configured small fast model, which defaults to Haiku. The model returns a yes-or-no decision and a short reason. A “no” tells Claude to keep working and includes the reason as guidance for the next turn. A “yes” clears the goal and records an achieved entry in the transcript.<p>> The evaluator runs on whichever provider your session is configured for. It does not call tools, so it can only judge what Claude has already surfaced in the conversation.<p>Apparently, it uses Haiku (by default) to evaluate every turn to determine if the goal has been achieved. However, it only relies on the transcript itself (including the reasoning of the main model). It can't independently verify if the goal has been achieved. So, if the main model thinks the goal is or isn't done, how often does Haiku disagree (in a productive way)? That's not clear to me.</p>
]]></description><pubDate>Tue, 23 Jun 2026 18:48:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48649559</link><dc:creator>kakugawa</dc:creator><comments>https://news.ycombinator.com/item?id=48649559</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48649559</guid></item></channel></rss>