<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: behohippy</title><link>https://news.ycombinator.com/user?id=behohippy</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 21 Jul 2026 21:12:27 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=behohippy" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by behohippy in "Kimi Work"]]></title><description><![CDATA[
<p>It's pretty simple nowadays if you know conceptually how they work.  Running the LLM calls in a loop with tools is an agent.  You only need 10 or so basic tools to accomplish nearly anything, and you can build a dynamic skill system from that.  Look at <a href="https://github.com/patw/pengy" rel="nofollow">https://github.com/patw/pengy</a>, ignore the app look at the spec.md file, feed that to your current agent of choice and make your own version.  Use whatever tech stack or UI you're comfortable with.  Change some of the choices in how it works, so it fits what you want to work.</p>
]]></description><pubDate>Tue, 21 Jul 2026 02:13:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48987279</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=48987279</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48987279</guid></item><item><title><![CDATA[New comment by behohippy in "ASML's Best Selling Product Isn't What You Think It Is"]]></title><description><![CDATA[
<p>You might have a business idea there.  I wouldn't mind a twinscan plushie for sitting on top of the workstation.</p>
]]></description><pubDate>Mon, 04 May 2026 11:32:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48007363</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=48007363</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48007363</guid></item><item><title><![CDATA[New comment by behohippy in "Benchmark Framework Desktop Mainboard and 4-node cluster"]]></title><description><![CDATA[
<p>Yeah 48g, sub 200W seems like a sweet spot for a single card setup.  Then you can stack as deep as you want to get the size of model you want for whatever you want to pay for the power bill.</p>
]]></description><pubDate>Mon, 11 Aug 2025 15:44:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=44865507</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=44865507</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44865507</guid></item><item><title><![CDATA[New comment by behohippy in "Ask HN: Do you think differently about working on open source these days?"]]></title><description><![CDATA[
<p>Sure, all the slop code projects I produce get MIT licensed on public repos.  It wasn't mine to begin with, so I wouldn't prevent anyone from using it.</p>
]]></description><pubDate>Mon, 11 Aug 2025 15:42:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=44865473</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=44865473</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44865473</guid></item><item><title><![CDATA[New comment by behohippy in "Benchmark Framework Desktop Mainboard and 4-node cluster"]]></title><description><![CDATA[
<p>Used 3090s have been getting expensive in some markets.  Another option is dual 5060ti 16 gig.  Mine are lower powered, single 8 pin power, so they max out around 180W.  With that I'm getting 80t/s on the new qwen 3 30b a3b models, and around 21t/s on Gemma 27b with vision.  Cheap and cheerful setup if you can find the cards at MSRP.</p>
]]></description><pubDate>Fri, 08 Aug 2025 01:08:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=44832309</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=44832309</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44832309</guid></item><item><title><![CDATA[New comment by behohippy in "Deepseek R1-0528"]]></title><description><![CDATA[
<p>About 768 gigs of ddr5 RAM in a dual socket server board with 12 channel memory and an extra 16 gig or better GPU for prompt processing.  It's a few grand just to run this thing at 8-10 tokens/s</p>
]]></description><pubDate>Wed, 28 May 2025 18:31:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=44119119</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=44119119</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44119119</guid></item><item><title><![CDATA[New comment by behohippy in "DeepSeek-V3 Technical Report"]]></title><description><![CDATA[
<p>These articles are gold, thank you.  I used your gemma one from a few weeks back to get gemma 3 performing properly.  I know you guys are all GPU but do you do any testing on CPU/GPU mixes? I'd like to see the pp and t/s on pure 12 channel epyc and the same with using a 24 gig gpu to accelerate the pp.</p>
]]></description><pubDate>Thu, 27 Mar 2025 11:37:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=43492541</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=43492541</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43492541</guid></item><item><title><![CDATA[New comment by behohippy in "Building a personal, private AI computer on a budget"]]></title><description><![CDATA[
<p>I run the KV cache at Q8 even on that model.  Is it not working well for you?</p>
]]></description><pubDate>Thu, 13 Feb 2025 18:37:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=43039504</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=43039504</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43039504</guid></item><item><title><![CDATA[New comment by behohippy in "Building a personal, private AI computer on a budget"]]></title><description><![CDATA[
<p>Qwen is a little fussy about the sampler settings, but it does run well quantized.  If you were getting infinite repetition loops, try dropping the top_p a bit.  I think qwen likes lower temps too</p>
]]></description><pubDate>Wed, 12 Feb 2025 10:49:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=43024026</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=43024026</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43024026</guid></item><item><title><![CDATA[New comment by behohippy in "Building a personal, private AI computer on a budget"]]></title><description><![CDATA[
<p>You probably won't be running fp16 anything locally.  We typically run Q5 or Q6 quants to maximize the size of the model and context length we can run with the VRAM we have available.  The quality loss is negligable at Q6.</p>
]]></description><pubDate>Tue, 11 Feb 2025 12:38:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=43012090</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=43012090</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43012090</guid></item><item><title><![CDATA[New comment by behohippy in "Ask HN: Is anyone doing anything cool with tiny language models?"]]></title><description><![CDATA[
<p>Just this pic:  <a href="https://imgur.com/ip8GWIh" rel="nofollow">https://imgur.com/ip8GWIh</a></p>
]]></description><pubDate>Wed, 22 Jan 2025 12:41:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=42792165</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=42792165</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42792165</guid></item><item><title><![CDATA[New comment by behohippy in "Ask HN: Is anyone doing anything cool with tiny language models?"]]></title><description><![CDATA[
<p>I don't have a video but here's a pic of the output:  <a href="https://imgur.com/ip8GWIh" rel="nofollow">https://imgur.com/ip8GWIh</a></p>
]]></description><pubDate>Wed, 22 Jan 2025 12:40:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=42792159</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=42792159</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42792159</guid></item><item><title><![CDATA[New comment by behohippy in "Ask HN: Is anyone doing anything cool with tiny language models?"]]></title><description><![CDATA[
<p>It's a 3b model so the creativity is pretty limited.  What helped for me was prompting for specific stories in specific styles.  I have a python script that randomizes the prompt and the writing style, including asking for specific author styles.</p>
]]></description><pubDate>Wed, 22 Jan 2025 12:40:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=42792152</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=42792152</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42792152</guid></item><item><title><![CDATA[New comment by behohippy in "Ask HN: Is anyone doing anything cool with tiny language models?"]]></title><description><![CDATA[
<p>I have a mini PC with an n100 CPU connected to a small 7" monitor sitting on my desk, under the regular PC.  I have llama 3b (q4) generating endless stories in different genres and styles.  It's fun to glance over at it and read whatever it's in the middle of making.  I gave llama.cpp one CPU core and it generates slow enough to just read at a normal pace, and the CPU fans don't go nuts.  Totally not productive or really useful but I like it.</p>
]]></description><pubDate>Tue, 21 Jan 2025 20:57:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=42785105</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=42785105</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42785105</guid></item><item><title><![CDATA[New comment by behohippy in "Phi-3 Technical Report"]]></title><description><![CDATA[
<p>I had this same issue with incomplete answers on longer summarization tasks.  If you ask it to "go on" it will produce a better completion, but I haven't seen this behaviour in any other model.</p>
]]></description><pubDate>Tue, 23 Apr 2024 15:35:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=40133129</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=40133129</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=40133129</guid></item><item><title><![CDATA[New comment by behohippy in "Microsoft Phi-2 model changes licence to MIT"]]></title><description><![CDATA[
<p>It's probably an evolution of the phi-1/1.5 "Textbooks are all you Need" training method:  <a href="https://arxiv.org/abs/2309.05463" rel="nofollow">https://arxiv.org/abs/2309.05463</a></p>
]]></description><pubDate>Sat, 06 Jan 2024 12:16:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=38890822</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=38890822</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38890822</guid></item><item><title><![CDATA[New comment by behohippy in "Amazon-Unveils-Q"]]></title><description><![CDATA[
<p>No joke, that would be an awesome LLM project name!</p>
]]></description><pubDate>Fri, 01 Dec 2023 15:24:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=38487772</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=38487772</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38487772</guid></item><item><title><![CDATA[New comment by behohippy in "Amazon-Unveils-Q"]]></title><description><![CDATA[
<p>Top_p and top_k are pretty important concepts for LLMs same as temperature so P,K,C and F are underutilized</p>
]]></description><pubDate>Thu, 30 Nov 2023 14:33:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=38473955</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=38473955</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38473955</guid></item><item><title><![CDATA[New comment by behohippy in "OpenLLaMA 13B Released"]]></title><description><![CDATA[
<p>Hey emad, thanks for SD and this!  What's the plan if Meta does Apache 2.0 for LLaMA?  Just keep going and making the 30b and 65b or build different models?</p>
]]></description><pubDate>Sun, 18 Jun 2023 18:04:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=36382660</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=36382660</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=36382660</guid></item><item><title><![CDATA[New comment by behohippy in "RedPajama 7B (an Apache 2.0-licensed LLaMa) is now available"]]></title><description><![CDATA[
<p>Vicuna-13b (4bit) got the answer right, the first time as well.</p>
]]></description><pubDate>Tue, 06 Jun 2023 19:04:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=36217573</link><dc:creator>behohippy</dc:creator><comments>https://news.ycombinator.com/item?id=36217573</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=36217573</guid></item></channel></rss>