<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: rshemet</title><link>https://news.ycombinator.com/user?id=rshemet</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 09 Oct 2026 02:30:46 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=rshemet" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by rshemet in "Whistle: Speech to Text in 16.9 MB"]]></title><description><![CDATA[
<p>hey, Roman here from Cactus, thank you for the feature!<p>opening this thread for questions/feedback if you have any</p>
]]></description><pubDate>Thu, 08 Oct 2026 19:34:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=50010876</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=50010876</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50010876</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots"]]></title><description><![CDATA[
<p>Did you give it a tool to increase temperature, or only one that sets temperature to an absolute value?<p>Either way, setting temperature to 5° is obviously wrong - even if it knew the current temperature - but models of this size can't reason about relative values very well.<p>Give it a tool to change temperature by a given amount, and see what happens!</p>
]]></description><pubDate>Tue, 11 Aug 2026 19:26:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49263229</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=49263229</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49263229</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots"]]></title><description><![CDATA[
<p>hey Kenny, Roman from Cactus here -<p>could you say more? What kind of home assistant / what stack</p>
]]></description><pubDate>Tue, 11 Aug 2026 04:28:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49253383</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=49253383</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49253383</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots"]]></title><description><![CDATA[
<p>it stands for Lets-not-be-sarcastic :)</p>
]]></description><pubDate>Tue, 11 Aug 2026 02:45:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49252767</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=49252767</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49252767</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots"]]></title><description><![CDATA[
<p>Roman from Cactus here -<p>yes you're right, there's only so much a 14MB model can do.<p>Needle excels at in-conext inference, with tightly defined environments. In our experience:<p>accurate descriptions + narrow tool scope = success</p>
]]></description><pubDate>Tue, 11 Aug 2026 02:39:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49252731</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=49252731</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49252731</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots"]]></title><description><![CDATA[
<p>there are android binaries you can ship in your own app - <a href="https://huggingface.co/Cactus-Compute/needle2/tree/main" rel="nofollow">https://huggingface.co/Cactus-Compute/needle2/tree/main</a><p>but if you're just looking for somewhere to try the model, use our in-browser playground! - <a href="https://cactuscompute.com/needle">https://cactuscompute.com/needle</a></p>
]]></description><pubDate>Tue, 11 Aug 2026 02:29:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49252657</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=49252657</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49252657</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots"]]></title><description><![CDATA[
<p>Hey! Roman here from Cactus - yes, we're putting putting together a detailed guide for ESP32.<p>In the meantime, if you have enough RAM for the current model (≈28MB), our repo will get you up & running:<p><a href="https://github.com/cactus-compute/needle" rel="nofollow">https://github.com/cactus-compute/needle</a></p>
]]></description><pubDate>Mon, 10 Aug 2026 22:48:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49250960</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=49250960</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49250960</guid></item><item><title><![CDATA[Show HN: Cactus v2 – On-device AI with cloud fallback]]></title><description><![CDATA[
<p>Hi HN, Roman and Henry here from Cactus (<a href="https://github.com/cactus-compute/cactus" rel="nofollow">https://github.com/cactus-compute/cactus</a>).<p>We just shipped the biggest upgrade to our on-device inference platform:<p>- Built-in model confidence-based routing to hand off inference runs to the cloud
- Converter for any PyTorch model
- Lossless 4-bit quantization (evals on our GitHub README)
- GPU acceleration on compatible devices (starting with Apple Metal)
- Minimal RAM footprint
- Runs on any Arm device: iOS, Android, Mac, DGX Spark, Raspberry Pi, and more<p>All in, a Gemma 4 E2B class model runs at 169 tok/sec on M5 Max, takes 2.7GB disk space with no accuracy degradation from FP16, uses 1.3GB of RAM, and requests help from cloud models when needed.<p>The problem we started with eighteen months ago: inference engines are built for datacenters, but consumer hardware has different physics: you share RAM with the OS, you get thermally throttled, and the same model behaves differently on different hardware.<p>So we wrote a runtime from scratch for resource-constrained devices. Since then, Cactus has grown to process millions of weekly inference runs, and tens of thousands of monthly active developers.<p>Our biggest learning from deploying Cactus in production apps is that while local models can handle 90% of workloads, that 10% gap means they're still not production-ready. Our users' fix was to build custom cloud fallback logic.<p>Cactus v2 fixes that:<p>Our approach to cloud fallback is to post-train a probe into the model's weights that reads its internal activations and emits a confidence signal. This way, the routing happens inside the model rather than in a prompt classifier sitting in front of it. We believe this is critical for multi-turn agentic work, where the model should know which turns are easy enough to be handled locally, and which are hard - and get handed off to the cloud. What ships today is single-turn routing for Gemma-4 E2B against a configurable escalation endpoint (Gemini, Claude, OpenAI-compatible, or your own endpoint).<p>Our next target is hybrid-native models for multi-turn agentic work. This is the genuinely unsolved problem. Our current probe-in-the-weights is showing promising results and we look to release the first model variants soon.<p>The hybrid variants are Gemma derivatives, released under the Gemma terms noted on the HF cards.<p>In addition to the hybrid routing, the runtime has SOTA quantization, which is lossless at 4bit, memory maps weights to decrease RAM footprint and runs cross-platform, with Python, Rust, React Native, Swift, and Kotlin bindings.<p>Disclaimer: Cactus is distributed as source-available - free for personal use and small companies; commercial license above that (Docker-style license).<p>You can get started on our GitHub: <a href="https://github.com/cactus-compute/cactus" rel="nofollow">https://github.com/cactus-compute/cactus</a> or by `brew install cactus-compute/cactus/cactus`.</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48864459">https://news.ycombinator.com/item?id=48864459</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Fri, 10 Jul 2026 19:57:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48864459</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=48864459</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48864459</guid></item><item><title><![CDATA[New comment by rshemet in "Launch HN: Cactus (YC S25) – AI inference on smartphones"]]></title><description><![CDATA[
<p>Yes! Cactus is optimized for mobile CPU inference and we're finishing internal testing of hybrid kernels that use the NPU, as well other chips.<p>We don't advise using GPUs on smartphones, since they're very energy-inefficient. Mobile GPU inference is actually the main driver behind the stereotype that "mobile inference drains your battery and heats up your phone".<p>Wrt to your last question – the short answer is yes, we'll have multimodal support. We currently support voice transcription and image understanding. We'll be expanding these capabilities to add more models, voice synthesis, and much more.</p>
]]></description><pubDate>Fri, 19 Sep 2025 10:46:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=45300133</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=45300133</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45300133</guid></item><item><title><![CDATA[New comment by rshemet in "Launch HN: Cactus (YC S25) – AI inference on smartphones"]]></title><description><![CDATA[
<p>indeed, this is exactly the goal! The license grants rights to commercial use, unlocks additional hardware acceleration, includes cloud telemetry, and offers significant savings over using cloud APIs.<p>In our deployments, we've seen open source models rival and even outperform lower-tier cloud counterparts. Happy to share some benchmarks if you like.<p>Our pricing is on a per-monthly-active-device basis, regardless of utilization. For voice-agent workflows, you typically hit savings as soon as you process over ≈2min of daily inference.</p>
]]></description><pubDate>Fri, 19 Sep 2025 10:37:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=45300081</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=45300081</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45300081</guid></item><item><title><![CDATA[New comment by rshemet in "Gemma 3 270M: Compact model for hyper-efficient AI"]]></title><description><![CDATA[
<p>you can run it in Cactus Chat (download from the Play Store)</p>
]]></description><pubDate>Fri, 15 Aug 2025 00:38:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=44907399</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44907399</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44907399</guid></item><item><title><![CDATA[New comment by rshemet in "Gemma 3 270M: Compact model for hyper-efficient AI"]]></title><description><![CDATA[
<p>you can also run it on Cactus - either in Cactus Chat from the App/Play Store or by using the Cactus framework to integrate it into your own app</p>
]]></description><pubDate>Fri, 15 Aug 2025 00:37:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=44907392</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44907392</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44907392</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: OWhisper – Ollama for realtime speech-to-text"]]></title><description><![CDATA[
<p>THIS IS THE BOMB!!! So excited for this one. Thanks for putting cool tech out there.</p>
]]></description><pubDate>Fri, 15 Aug 2025 00:26:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=44907329</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44907329</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44907329</guid></item><item><title><![CDATA[New comment by rshemet in "I want everything local – Building my offline AI workspace"]]></title><description><![CDATA[
<p>if you ever end up trying to take this in the mobile direction, consider running on-device AI with Cactus –<p><a href="https://cactuscompute.com/">https://cactuscompute.com/</a><p>Blazing-fast, cross-platform, and supports nearly all recent OS models.</p>
]]></description><pubDate>Fri, 08 Aug 2025 19:30:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=44840769</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44840769</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44840769</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Cactus – Ollama for Smartphones"]]></title><description><![CDATA[
<p><a href="https://play.google.com/store/apps/details?id=com.rshemetsubuser.myapp">https://play.google.com/store/apps/details?id=com.rshemetsub...</a></p>
]]></description><pubDate>Fri, 11 Jul 2025 16:37:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=44534214</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44534214</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44534214</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Cactus – Ollama for Smartphones"]]></title><description><![CDATA[
<p>thank you! Very kind feedback, and we'll add your feedback to our to-dos.<p>re: "question would get stuck on the last phrase and keep repeating it without end." - that's a limitation of the model i'm afraid. Smaller models tend to do that sometimes.</p>
]]></description><pubDate>Fri, 11 Jul 2025 16:36:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=44534199</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44534199</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44534199</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Cactus – Ollama for Smartphones"]]></title><description><![CDATA[
<p>say more about "community tools"?</p>
]]></description><pubDate>Fri, 11 Jul 2025 16:35:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=44534187</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44534187</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44534187</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Cactus – Ollama for Smartphones"]]></title><description><![CDATA[
<p>in the app you mean?<p>Adding shortly!</p>
]]></description><pubDate>Fri, 11 Jul 2025 16:34:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=44534184</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44534184</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44534184</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Cactus – Ollama for Smartphones"]]></title><description><![CDATA[
<p>that's our mission! if you are passionate about the space, we look forward to your contributions!</p>
]]></description><pubDate>Fri, 11 Jul 2025 16:34:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=44534177</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44534177</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44534177</guid></item><item><title><![CDATA[New comment by rshemet in "Show HN: Cactus – Ollama for Smartphones"]]></title><description><![CDATA[
<p>no, good observation - not hidden; we don't have a "clear conversation" button.<p>to your previous point - Cactus fully supports tool calling (for models that have been instruction-trained accordingly, e.g. Qwen 1.7B)<p>for "turning your old phones into local llm servers", Cactus is likely not the best tool. We'd recommend something like actual Ollama or Exo</p>
]]></description><pubDate>Fri, 11 Jul 2025 16:33:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=44534168</link><dc:creator>rshemet</dc:creator><comments>https://news.ycombinator.com/item?id=44534168</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44534168</guid></item></channel></rss>