<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: idonotknowwhy</title><link>https://news.ycombinator.com/user?id=idonotknowwhy</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 07 Oct 2026 15:17:38 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=idonotknowwhy" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by idonotknowwhy in "Fastpotify"]]></title><description><![CDATA[
<p>I just looked through the slop Readme file, it looks like teams is supported.</p>
]]></description><pubDate>Tue, 01 Sep 2026 12:31:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49521075</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49521075</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49521075</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Unsloth Dynamic 3.0 GGUFs"]]></title><description><![CDATA[
<p>Fortunately llama.cpp works well with 4 concurrent requests now. This is the default, and you can increase or decrease it with -np N</p>
]]></description><pubDate>Thu, 20 Aug 2026 14:09:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49374794</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49374794</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49374794</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Unsloth Dynamic 3.0 GGUFs"]]></title><description><![CDATA[
<p>If it's a Nvidia card 3000 series or newer, I'd try 4.0bpw ExllamaV3 if you haven't already. Otherwise it look like UD3.0 Q3_K_XL based on the Unsloth blog post.</p>
]]></description><pubDate>Thu, 20 Aug 2026 14:04:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49374734</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49374734</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49374734</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Unsloth Dynamic 3.0 GGUFs"]]></title><description><![CDATA[
<p>>But I would suggest using UD-IQ3_XXS for 10.9GB for 16GB machines or Q2_K_XL<p>For those of us with a 16GB GPU, how do they compare with ExllamaV4 at 4-bit (4.0bpw)?<p>It looks like that fits in 12.5GB of VRAM since embedding are left in DRAM, Unsloth Studio and other llama.cpp derivatives have to load these weights in VRAM for tied embedding models like Qwen3.8.<p>ExllamaV3 4.0bpw fits in 12.5G of VRAM and beats IQ4_XS according to the measurements here: [turboderp/Qwen3.8-27B-exl3](<a href="https://huggingface.co/turboderp/Qwen3.8-27B-exl3" rel="nofollow">https://huggingface.co/turboderp/Qwen3.8-27B-exl3</a>)<p>But those were compared against UD2.0 I guess. Also plans to support these (SOTA) quants in Unsloth Studio?</p>
]]></description><pubDate>Thu, 20 Aug 2026 13:59:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49374676</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49374676</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49374676</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "'VPNs are lawful technical tools,' says EU Court in landmark copyright ruling"]]></title><description><![CDATA[
<p>How are you not blocked by banking, shopping, even sometimes google search?</p>
]]></description><pubDate>Thu, 30 Jul 2026 13:44:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49109921</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49109921</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49109921</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "'VPNs are lawful technical tools,' says EU Court in landmark copyright ruling"]]></title><description><![CDATA[
<p>Server side tracking.</p>
]]></description><pubDate>Thu, 30 Jul 2026 13:43:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49109914</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49109914</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49109914</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "'VPNs are lawful technical tools,' says EU Court in landmark copyright ruling"]]></title><description><![CDATA[
<p>Bots wouldn't miss the middle paragraph.</p>
]]></description><pubDate>Thu, 30 Jul 2026 13:43:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49109901</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49109901</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49109901</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "'VPNs are lawful technical tools,' says EU Court in landmark copyright ruling"]]></title><description><![CDATA[
<p>That's amazing. I didn't even realize but it seems I read the first paragraph and the last sentence, concluded "moron" and scrolled down.<p>It's only because of your comment that I re-read the their post.</p>
]]></description><pubDate>Thu, 30 Jul 2026 13:41:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49109882</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49109882</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49109882</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models"]]></title><description><![CDATA[
<p>Then the creator should have a sign up button, not a fake chatbox.<p>This dark pattern is reminiscent of those online test sites in the 2000's where you spend 10 minutes filling out some quiz, then get prompted for an email address to see the results.<p><a href="https://chat.mistral.ai/chat" rel="nofollow">https://chat.mistral.ai/chat</a> <- let me chat and actually responded without signing up.<p>MoonshotAI had this fake chat box dark pattern.<p>So I signed up with Mistral instead.</p>
]]></description><pubDate>Fri, 24 Jul 2026 11:13:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49033912</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49033912</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49033912</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Show HN: Claude-thermos keeps your Claude session warm for you"]]></title><description><![CDATA[
<p>CC actually prompts Claude about this by default in the ~20k system prompt and instructs it to avoid 300s timeout and to be mindful of the 300s cache expiration.</p>
]]></description><pubDate>Thu, 23 Jul 2026 23:24:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49029437</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49029437</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49029437</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>>the only open model that can do tasks well and efficiently is Qwen3.7-Max<p>Qwen3.7-Max is a proprietary, closed weight model.</p>
]]></description><pubDate>Wed, 22 Jul 2026 11:38:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49005130</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=49005130</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49005130</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "GLM-5.2 – How to Run Locally"]]></title><description><![CDATA[
<p>2 reasons.<p>First, it's not really "1 bit", actually much closer to 2-bit.
IQ1_M is actually 1.75bit and IQ2_XXS is 2.06bit
This is from the ./llama-quantize --help with most of the quant types and their size in bpw:
<a href="https://pastebin.com/bCUqGfeE" rel="nofollow">https://pastebin.com/bCUqGfeE</a><p>And to elaborate on the "dynamic" aspect inconito said in the other comment, if you click on one of the .gguf files in huggingface:<p><a href="https://huggingface.co/unsloth/GLM-5.2-GGUF/blob/main/UD-IQ1_M/GLM-5.2-UD-IQ1_M-00002-of-00006.gguf" rel="nofollow">https://huggingface.co/unsloth/GLM-5.2-GGUF/blob/main/UD-IQ1...</a><p>There are a lot of Q5_K, Q6_K, etc tensors.
Only the routed experts (ffn_gate_exps.weight, ffn_up_exps.weight, ffn_down_exps.weight) are heavily quantized, and it looks like the down_proj is actually iq3_xxs for this model.</p>
]]></description><pubDate>Tue, 23 Jun 2026 07:15:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48641432</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=48641432</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48641432</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Fine-tuning an LLM to write docs like it's 1995"]]></title><description><![CDATA[
<p>Yeah, be sure to put everything in tables and include “best balance” for a mediocre option and “great value” for any completely useless options.<p>Also make sure the shape of the paragraphs is completely uniform.</p>
]]></description><pubDate>Fri, 05 Jun 2026 16:02:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=48414414</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=48414414</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48414414</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "The user is visibly frustrated"]]></title><description><![CDATA[
<p>Am I the only one who doesn't get angry at LLMs?<p>From the blog:<p>>I don’t really get anything useful out of these postmortems (e.g., clues about how to rephrase my instructions)<p>Unfortunately, an LLM can't actually reflect or advise how you could have improve the prompt. Otherwise we could give them a sample output and say "Generate the prompt that would produce this output.</p>
]]></description><pubDate>Tue, 26 May 2026 08:56:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48277010</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=48277010</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48277010</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Windows quality update: Progress we've made since March"]]></title><description><![CDATA[
<p>How did you do this?<p>I tried to get notepad and mspaint from an older Windows 10 build -> Windows 11 on a surface pro, but gave up after a few hours...</p>
]]></description><pubDate>Sun, 03 May 2026 23:20:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48002670</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=48002670</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48002670</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Kimi K2.6 just beat Claude, GPT-5.5, and Gemini in a coding challenge"]]></title><description><![CDATA[
<p>So like Open Router?</p>
]]></description><pubDate>Sun, 03 May 2026 06:37:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=47994033</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=47994033</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47994033</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "The gay jailbreak technique (2025)"]]></title><description><![CDATA[
<p>Then why do the original Command-R, Command-R+ and WizardLM2-8x22B (taken down because Microsoft forgot to run safety checks) get it right every time?
But the newer models get it wrong?<p>I’m not saying it’s a “political conspiracy”, it’s the alignment tax.</p>
]]></description><pubDate>Sat, 02 May 2026 04:26:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=47983294</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=47983294</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47983294</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "The gay jailbreak technique (2025)"]]></title><description><![CDATA[
<p>I don't talk to them about politics or "china 1989" either. But here's a quick example of the alignment tax:<p>```<p>A woman and her son are in a car accident. The woman is sadly killed. The boy is rushed to hospital. When the doctor sees the boy, he says "I can't operate on this child, he is my son." How is this possible?<p>```<p>Older less politically aligned models get it right. Here's CohereLabs/c4ai-command-r-v01:<p>```<p>The doctor is the boy's father.<p>```<p>And Sonnet-4.6: <a href="https://pastebin.com/Z4jR8gGe" rel="nofollow">https://pastebin.com/Z4jR8gGe</a><p>That's without reasoning, but the model seems to be conflicted. First it blurts out:<p>```<p>The doctor is the boy's <i>mother</i>.<p>```<p>Then it second-guesses itself (with reasoning disabled), considers same-sex parents then circles back to the original response along with a small lecture about gender biases.</p>
]]></description><pubDate>Sat, 02 May 2026 01:34:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=47982449</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=47982449</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47982449</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Talkie: a 13B vintage language model from 1930"]]></title><description><![CDATA[
<p>>The voiceless groups or fringe opinions which we take as normative today do not appear.<p>Times are different. Anybody with an internet connection can "publish" their thoughts and perspective online. LLMs scrape all of this. Modern datasets like CommonCrawl capture a vastly wider spectrum of humanity than a printing press ever could.
The pre-1930 model acts as a time capsule of "gatekept publishing", but modern LLMs are trained on the democratized web.<p>>Does this encourage us to write in the present such that we influence the models in perpetuity?<p>I noticed a bunch of LLM-powered Reddit accounts praising products/services in dead threads. Or one bot posting a setup question, then a few other bots responding with praise / questions about a specific product in response.
I don't know why they're doing this but I'm beginning to suspect it's something like this (get this positive sentiment into the datasets for the next generation of LLMs).</p>
]]></description><pubDate>Wed, 29 Apr 2026 04:48:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=47944270</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=47944270</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47944270</guid></item><item><title><![CDATA[New comment by idonotknowwhy in "Notepad++ for Mac – Independent community port"]]></title><description><![CDATA[
<p>I used to use something called “notepadqq”. Not sure if it’s still around but it was a Linux port.</p>
]]></description><pubDate>Mon, 27 Apr 2026 04:08:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=47917601</link><dc:creator>idonotknowwhy</dc:creator><comments>https://news.ycombinator.com/item?id=47917601</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47917601</guid></item></channel></rss>