<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: waximabbax</title><link>https://news.ycombinator.com/user?id=waximabbax</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 10 Oct 2026 09:11:39 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=waximabbax" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by waximabbax in "Claude Haiku 5.5"]]></title><description><![CDATA[
<p>Alright its still little early since there is not enough independent testing but this looks very promising and I wasn't expecting anthropic to beat GPT-6 Luna especially at the same price. Haiku 5.5 beats Luna on every shared benchmark Anthropic published, particularly computer use and agentic coding.</p>
]]></description><pubDate>Wed, 07 Oct 2026 20:13:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49998170</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49998170</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49998170</guid></item><item><title><![CDATA[New comment by waximabbax in "Show HN: Open-source model routing for coding agents at Astra-level performance"]]></title><description><![CDATA[
<p>It mostly feels horribly slow if you leave reasoning effort at max for Sol 6.1. If you dial it to normal/high/xhigh, Sol 6.1 is only a couple of minutes behind Astra, performs almost as well on terminal bench v4 tasks, and is about 4 to 5x cheaper.</p>
]]></description><pubDate>Fri, 02 Oct 2026 17:09:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49935926</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49935926</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49935926</guid></item><item><title><![CDATA[New comment by waximabbax in "Typed-lm: a Rust jev open source alternative"]]></title><description><![CDATA[
<p>Its use case is still pretty narrow, just like Jev. The only time Jev could make sense is if you want thousands of requests per second and can sacrifice a bit of accuracy. Otherwise, modern LLMs are super cheap, like GPT-6 Luna, GLM-5.3-Flash, or even Gemma 4, and reason much better than Jev with better accuracy while being perfectly capable of returning structured JSON. So it’s basically enforcing typed answers with lower accuracy vs. occasionally getting the response format wrong but having better accuracy.</p>
]]></description><pubDate>Sat, 26 Sep 2026 12:51:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49856030</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49856030</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49856030</guid></item><item><title><![CDATA[New comment by waximabbax in "Anthropic: The Situation Report"]]></title><description><![CDATA[
<p>Probably one of the better real-world AI use cases I’ve seen. Automate the tedious data work, save hours, and keep the actual scientific decisions with experts.</p>
]]></description><pubDate>Fri, 25 Sep 2026 11:51:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49843311</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49843311</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49843311</guid></item><item><title><![CDATA[New comment by waximabbax in "Repo indexing was doing more harm than good, so we dropped it from our agent"]]></title><description><![CDATA[
<p>A few years ago, when terminal agents were still somewhat new, indexing was thought to be the holy grail for coding agents. But in most cases, it did us more harm than good. More often than not, it confused the model instead of letting it directly grep for context and work its way around the codebase. Not to mention that it also increased input token costs and added extra round trips.<p>Similarity search is rarely that relevant when working on codebases. Our biggest learning actually came from an accidental technical bug where indexing had been returning zero hits for quite some time, and we barely noticed any performance degradation. In fact, the agent seemed to be working better than ever.<p>After discovering the bug, we did some rigorous A/B testing and realized it was best to drop indexing altogether. There could still be a case for it in extremely large codebases with a lot of docs e.g. excel or several text files.</p>
]]></description><pubDate>Fri, 28 Aug 2026 09:56:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49476515</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49476515</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49476515</guid></item><item><title><![CDATA[Repo indexing was doing more harm than good, so we dropped it from our agent]]></title><description><![CDATA[
<p>Article URL: <a href="https://thegit.ai/blog/why-we-abandoned-repository-indexing">https://thegit.ai/blog/why-we-abandoned-repository-indexing</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49476449">https://news.ycombinator.com/item?id=49476449</a></p>
<p>Points: 1</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 28 Aug 2026 09:47:12 +0000</pubDate><link>https://thegit.ai/blog/why-we-abandoned-repository-indexing</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49476449</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49476449</guid></item><item><title><![CDATA[New comment by waximabbax in "RAG Is Simpler Than You Think"]]></title><description><![CDATA[
<p>We removed retrieval from our coding agent a while back. What convinced us wasn’t a benchmark, we found that the retrieval path had been returning zero results for quite some time because of a technical bug, still nobody noticed, indeed it was working better than before.<p>After doing some rigorous A/B testing, we dropped indexing. For coding, I think the reason is that a repo is already searchable. Imports, call sites, file and test names, grep gives you cheap yet reliable version of what indexing would do, and the agent can read around a hit to verify it. Chunked retrieval hands the model something that looks right, and it tends to trust that instead of going to look for the actual source. Another thing that I noticed was the most intelligent models like Opus 5 and Fable ignored chunks anyway most of the time for some reason. Possibly perhaps they are trained around not trusting similarity checks for codebases.<p>Extremely large codebases with docs feel different. You can’t grep for a concept you can’t name. That’s the case where I’d still use retrieval.<p>(I work on TheGitAI, for disclosure.)</p>
]]></description><pubDate>Wed, 26 Aug 2026 15:47:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49451182</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49451182</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49451182</guid></item><item><title><![CDATA[Show HN: TheGitAI – A model-agnostic coding agent for your terminal]]></title><description><![CDATA[
<p>Article URL: <a href="https://thegit.ai/">https://thegit.ai/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49436326">https://news.ycombinator.com/item?id=49436326</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 25 Aug 2026 16:03:11 +0000</pubDate><link>https://thegit.ai/</link><dc:creator>waximabbax</dc:creator><comments>https://news.ycombinator.com/item?id=49436326</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49436326</guid></item></channel></rss>