<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: throwdbaaway</title><link>https://news.ycombinator.com/user?id=throwdbaaway</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 01 Aug 2026 00:27:35 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=throwdbaaway" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by throwdbaaway in "DeepSeek-V4-Flash Update"]]></title><description><![CDATA[
<p>Objectively speaking, the 2 bit quant from antirez has very low accuracy. Meanwhile, his 4 bit quant does have decent accuracy, but is a bit pointless by being bigger than the full precision MXFP4 quant. Anyway, they all work fine in practice.</p>
]]></description><pubDate>Fri, 31 Jul 2026 11:49:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49121960</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=49121960</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49121960</guid></item><item><title><![CDATA[New comment by throwdbaaway in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>They need to get a license from moonshot to provide inference for K3. Probably have to follow the pricing set by moonshot as well.</p>
]]></description><pubDate>Mon, 27 Jul 2026 15:31:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49071074</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=49071074</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49071074</guid></item><item><title><![CDATA[New comment by throwdbaaway in "Laguna S 2.1"]]></title><description><![CDATA[
<p>It works, thanks to <a href="https://github.com/ikawrakow/ik_llama.cpp/pull/1911" rel="nofollow">https://github.com/ikawrakow/ik_llama.cpp/pull/1911</a>, which got merged in early June. However, there might still be some issue with the chat template.</p>
]]></description><pubDate>Wed, 22 Jul 2026 08:56:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49003694</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=49003694</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49003694</guid></item><item><title><![CDATA[New comment by throwdbaaway in "Qwen 3.8"]]></title><description><![CDATA[
<p>I suspect this is why DeepSeek had to introduce the 2x peak hours pricing. The price would be too low otherwise.</p>
]]></description><pubDate>Sun, 19 Jul 2026 23:09:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48972535</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48972535</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48972535</guid></item><item><title><![CDATA[New comment by throwdbaaway in "Control the Ideas, Not the Code"]]></title><description><![CDATA[
<p>Yeah antirez made a lot of big claims in that paragraph. Sounds like a case of AI psychosis.</p>
]]></description><pubDate>Tue, 14 Jul 2026 20:03:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48912288</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48912288</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48912288</guid></item><item><title><![CDATA[New comment by throwdbaaway in "Show HN: Getting GLM 5.2 running on my slow computer"]]></title><description><![CDATA[
<p>If you max out the ram, TG with q3 should be at least 10 t/s. And with dsa, it can still stay close to that number as the context grows.</p>
]]></description><pubDate>Fri, 10 Jul 2026 16:24:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=48862081</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48862081</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48862081</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>That's exactly what I said. They do care when FLOPs are involved. Restoring an old session with 900k tokens will require a lot of FLOPs to reprocess the 900k token.<p>Meanwhile, they don't really care if you use hundreds of millions of cached input tokens, which doesn't consume any FLOP.</p>
]]></description><pubDate>Tue, 07 Jul 2026 16:49:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48820367</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48820367</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48820367</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>Different sessions. With <a href="https://github.com/fairydreaming/llama.cpp/tree/dsv4" rel="nofollow">https://github.com/fairydreaming/llama.cpp/tree/dsv4</a>, 1M context with DSV4 Flash takes less than 6GB of VRAM. I can't run DSV4 Pro, but it should take less than 9GB of VRAM for 1M context, based on the numbers shared in <a href="https://arxiv.org/html/2606.19348v1" rel="nofollow">https://arxiv.org/html/2606.19348v1</a>.</p>
]]></description><pubDate>Tue, 07 Jul 2026 07:20:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48814659</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48814659</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48814659</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>Well I wouldn't call it a low bar, since some of the edits were quite complex. And 1M context in less than 6GB of VRAM is truly impressive, but somehow this gets way less attention than the crappy turbo quant from Google.</p>
]]></description><pubDate>Tue, 07 Jul 2026 05:47:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=48814069</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48814069</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48814069</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>While we are all speculating, Boris kindly provided some guidance in <a href="https://news.ycombinator.com/item?id=47880089">https://news.ycombinator.com/item?id=47880089</a><p>> The challenge is: when you let a session idle for >1 hour, when you come back to it and send a prompt, it will be a full cache miss, all N messages. We noticed that this corner case led to outsized token costs for users. In an extreme case, if you had 900k tokens in your context window, then idled for an hour, then sent a message, that would be >900k tokens written to cache all at once, which would eat up a significant % of your rate limits, especially for Pro users.<p>Using the current Opus pricing, that pre-lunch 900k tokens should roughly consist of:<p>720k input tokens = 0.72 x $5 = $3.6<p>180k output tokens = 0.18 x $25 = $4.5<p>900k 1h cached writes = 0.9 x $10 = $9<p>500M cached input tokens = 500 x $0.5 = $250<p>$267.1 in total, with 93.6% from cached input tokens. The portion that requires GPU compute is about 3% of the total.<p>Post-lunch, the 900k tokens should consist of:<p>900k input tokens = 0.9 x $5 = $4.5<p>900k 1h cached writes = 0.9 x $10 = $9<p>So Anthropic is fine with the $267.1 accumulated over 3~4 hours before lunch, but not fine with the $13.5 incurred immediately after lunch. Why?<p>The only plausible explanation is that the actual cost of caching is way less than the API pricing. If you use a coding plan, Anthropic doesn't really care about your cached input tokens usage. Indeed they want you to show your ccusage screenshots. On the other hand, if you pay by API tokens, the margin is huge for cached input tokens.<p>Only when you do something that requires a lot of FLOPs, e.g. the post-lunch 900k input tokens, the cost becomes real.</p>
]]></description><pubDate>Mon, 06 Jul 2026 23:49:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=48811988</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48811988</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48811988</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>Indeed they are all lossy. Not sure how much they contribute to the quality loss in long context though. I got a 700k session with DSV4 Pro (official API), and the model was still coherent and didn't make any tool call error.</p>
]]></description><pubDate>Mon, 06 Jul 2026 23:29:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48811846</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48811846</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48811846</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>The current top comment in <a href="https://lobste.rs/s/ua1gxl/glm_5_2_coming_ai_margin_collapse" rel="nofollow">https://lobste.rs/s/ua1gxl/glm_5_2_coming_ai_margin_collapse</a> correctly zoomed into cached input tokens, but landed on the opposite conclusion:<p>> That is, for your $100/month fee, you get $3600 equivalent of API usage. This is presumably because Anthropic has figured out some clever things to do with model routing and input caching, and also can subsidize with investor money and take a hit on their operating margins.<p>My take: this is exactly what Anthropic wants everyone to think. In reality, 90% of that $3600 are for cached input tokens, that can be made to cost next to nothing, as shown by DeepSeek.</p>
]]></description><pubDate>Mon, 06 Jul 2026 23:05:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48811635</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48811635</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48811635</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>Seems like a pretty pointless post that still centers around output tokens.<p>In agentic coding, cached input tokens is 90% of the API "cost". It doesn't require GPU compute, and DeepSeek has shown that it can be done 50~100x cheaper with MLA/CSA/HCA, and a whole bunch of disks. This should collapse the margin.</p>
]]></description><pubDate>Mon, 06 Jul 2026 22:52:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48811532</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48811532</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48811532</guid></item><item><title><![CDATA[New comment by throwdbaaway in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>And somehow they claimed that it is "lossless".</p>
]]></description><pubDate>Sat, 04 Jul 2026 01:10:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48781759</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48781759</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48781759</guid></item><item><title><![CDATA[New comment by throwdbaaway in "GLM-5.2 – How to Run Locally"]]></title><description><![CDATA[
<p>On ZFS with zstd compression, I am getting 1.34x compressratio for the BF16 weights (across multiple models).<p>Here's the du output for GLM-5.2:<p><pre><code>    $ du -s -BG /cube/models/zai-org/GLM-5.2/
    1099G   /cube/models/zai-org/GLM-5.2/</code></pre></p>
]]></description><pubDate>Tue, 23 Jun 2026 03:51:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=48640037</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48640037</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48640037</guid></item><item><title><![CDATA[New comment by throwdbaaway in "DeepSeek makes the V4 Pro price discount permanent"]]></title><description><![CDATA[
<p>And their disk-based caching is amazing. I got a long 700k context session spanning more than a week, with pauses in between that was longer than a day, and some rewinds mixed in as well.<p>Stats from pi:<p>↑400k ↓438k R432M 71.9%/1.0M<p>Half a billion tokens, $2.12</p>
]]></description><pubDate>Sat, 23 May 2026 01:28:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=48243628</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48243628</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48243628</guid></item><item><title><![CDATA[New comment by throwdbaaway in "A few words on DS4"]]></title><description><![CDATA[
<p>Hah, that's because the prompt itself was only about 30 tokens. We need a much bigger prompt to properly test PP.</p>
]]></description><pubDate>Fri, 15 May 2026 07:04:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48145419</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=48145419</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48145419</guid></item><item><title><![CDATA[New comment by throwdbaaway in "An AI agent deleted our production database. The agent's confession is below"]]></title><description><![CDATA[
<p>Huh that's not what I gathered from the tweet at all. If I am going to write a five why's analysis, the immediate cause is the LLM wrongly decided to delete a volume, while the root cause is the bad design to co-locate staging and production data in the same volume. The writing was quite vague though, let's wait for a response from railway.</p>
]]></description><pubDate>Mon, 27 Apr 2026 01:23:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=47916712</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=47916712</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47916712</guid></item><item><title><![CDATA[New comment by throwdbaaway in "An AI agent deleted our production database. The agent's confession is below"]]></title><description><![CDATA[
<p>If I understand correctly, both the staging database and the production database share the same volume. Thus, production data was gone as well after deleting the volume.<p>1st hint - the API call only contains one volume:<p><pre><code>    curl -X POST https://backboard.railway.app/graphql/v2 \
      -H "Authorization: Bearer [token]" \
      -d '{"query":"mutation { volumeDelete(volumeId: \"3d2c42fb-...\") }"}'
</code></pre>
2nd hint - this gem from the tweet:<p>> No "this volume contains production data, are you sure?"</p>
]]></description><pubDate>Sun, 26 Apr 2026 21:57:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=47915102</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=47915102</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47915102</guid></item><item><title><![CDATA[New comment by throwdbaaway in "An update on recent Claude Code quality reports"]]></title><description><![CDATA[
<p>Should be about 10~20 GiB per session. Save/restore is exactly what DeepSeek does using its 3FS distributed filesystem: <a href="https://github.com/deepseek-ai/3fs#3-kvcache" rel="nofollow">https://github.com/deepseek-ai/3fs#3-kvcache</a><p>With this much cheaper setup backed by disks, they can offer much better caching experience:<p>> Cache construction takes seconds. Once the cache is no longer in use, it will be automatically cleared, usually within a few hours to a few days.</p>
]]></description><pubDate>Thu, 23 Apr 2026 21:54:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=47882623</link><dc:creator>throwdbaaway</dc:creator><comments>https://news.ycombinator.com/item?id=47882623</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47882623</guid></item></channel></rss>