<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: HenryNdubuaku</title><link>https://news.ycombinator.com/user?id=HenryNdubuaku</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 20 Sep 2026 12:01:08 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=HenryNdubuaku" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>thanks, we improved on False Negatives this time :)</p>
]]></description><pubDate>Sat, 19 Sep 2026 05:16:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49763541</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49763541</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49763541</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>noted, we'd look into this, thanks</p>
]]></description><pubDate>Sat, 19 Sep 2026 00:07:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49761924</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49761924</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49761924</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>thanks!, let us know if you ever build it out :)</p>
]]></description><pubDate>Sat, 19 Sep 2026 00:07:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49761919</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49761919</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49761919</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thanks for testing Needle out! I'd be very interested in hearing more about the finetuning setup to see how we can make both the library's finetuning setup and the model better.</p>
]]></description><pubDate>Fri, 18 Sep 2026 21:35:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760586</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760586</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760586</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thanks for considering needle. Keep in mind that you can also fine-tune the model to fit your use case more. I think this illustrates the intended deployment pretty well, where both computational resources and compute credits can both be issues for deployment.</p>
]]></description><pubDate>Fri, 18 Sep 2026 21:31:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760535</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760535</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760535</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thank you fr this. "warm the house" now goes to the thermostat. It's fair that a more deterministic system with just action phrases would be easier to debug/interpret, but I think there is room for both a model that is trained to understand meaning as well as deterministic logic aiding it. To this end, we just started exploring the idea of triggers, and are working towards expanding this even more.</p>
]]></description><pubDate>Fri, 18 Sep 2026 21:29:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760498</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760498</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760498</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>I think we might look into creating a baseline like this for our future models</p>
]]></description><pubDate>Fri, 18 Sep 2026 21:25:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760456</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760456</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760456</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thanks!</p>
]]></description><pubDate>Fri, 18 Sep 2026 21:09:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760286</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760286</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760286</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>The model can be sliced and perform the inference using a subset of its layers. The first 4 layers alone are 8 MB, all 20 are 29 MB. Fair point on the copy, tightened it.</p>
]]></description><pubDate>Fri, 18 Sep 2026 21:03:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760234</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760234</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760234</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thank you!</p>
]]></description><pubDate>Fri, 18 Sep 2026 20:54:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49760129</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49760129</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49760129</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>I think that would be very useful for us! The best way to reach us is through the founders@cactuscompute.com email<p>Thank you!</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:54:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49759366</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49759366</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49759366</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>This is extremely useful feedback for us, thanks! I think the easiest thing here that can be fixed with tool definitions is the number conversions. Additionally, the model tends to work better with fewer tools. We will definitely be focusing on better context usage and followups going forward as well.</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:53:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49759337</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49759337</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49759337</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thanks for the feedback! Implications and relations are hard for the model to understand (things like go to the living room, then the kitchen, and back), so yes the cleanest use cases involve direct language. Reasoning isn't true <i>reasoning</i> in the way general LLMs do it, it is more like grounding for the model that it generates itself. This can often become nonsensical specifically when the model gets things wrong, providing signal to the confidence.</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:48:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49759277</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49759277</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49759277</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>That's a really good point and I think it's not yet clear how well, say, 8-30MB worth of regexs with accompanying algorithmic structure would do on these tasks. I would imagine they do quite well on a well defined task, but it would be much harder to then adapt this set to a new domain. A big part of Needle's promise is how easy it is to finetune. Ultimately I think the two approaches can be more complimentary to each other, rather than choosing only one (see triggers!).</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:30:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49759061</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49759061</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49759061</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Thanks! Certainly giving the model more context on the task it needs to perform would help it. This was actually a part of training that we improved going from Needle 2 to Needle 3</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:19:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49758934</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49758934</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49758934</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>As far as I know Jev's architecture isn't public (though I might be mistaken!), but it is a coincidence :)<p>Needle 3 has been in the making since Needle 2 launched early august, but we are very excited that Jev is bringing more attention to the problem we are trying to solve.</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:18:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49758921</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49758921</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49758921</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>well certainly the environment on the website cannot be a full product, and it isn't claiming to be that. The model, while capable in many dimensions, is also limited by its size. The website is meant to show both the capabilities and the limitations! A real deployment would absolutely need external guardrails, more thoroughly thought out tool sets with better task-specific triggers, perhaps also task-specific finetuning for better confidence grounding. And in my view that's the point of small open models! You can take it and run with it as far as you want.</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:16:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49758900</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49758900</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49758900</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Hey there! If you end up trying out needle on the test suite it would be very useful for us if you could share some failure modes of the model! We are always trying to understand where the model isn't doing good and where we can make it better.<p>For your question on tool calling, I think you will find that the model is pretty good at simpler tool calls and parallel ones, but can struggle with implied references and multistep reasoning. These are definitely things that can improve with task-specific finetuning but for some things you just have to have a model that is properly sized. That said, we are always trying to improve the model so that it can handle an ever larger set of queries</p>
]]></description><pubDate>Fri, 18 Sep 2026 19:00:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49758692</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49758692</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49758692</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Hey, thanks for the feedback! I think this is a useful part of a demonstration so I added a 911 tool specifically to demonstrate this capability and the fact that you can guard it with triggers that make it so calling emergency is an unambiguous action given the input. This really shows that constructing the right tool set with the right surrounding setup is a priority when deploying needle.</p>
]]></description><pubDate>Fri, 18 Sep 2026 18:54:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=49758608</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49758608</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49758608</guid></item><item><title><![CDATA[New comment by HenryNdubuaku in "Show HN: Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash"]]></title><description><![CDATA[
<p>Hey there, yep we found that on Apple devices specifically running on CPU is fast enough that Metal support is not needed. Thanks for flagging this though, and if usecases that would benefit from Metal support come up we will be adding it to the binaries.</p>
]]></description><pubDate>Fri, 18 Sep 2026 18:51:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49758578</link><dc:creator>HenryNdubuaku</dc:creator><comments>https://news.ycombinator.com/item?id=49758578</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49758578</guid></item></channel></rss>