<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: WASDx</title><link>https://news.ycombinator.com/user?id=WASDx</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 13 Aug 2026 21:38:11 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=WASDx" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by WASDx in "DeepSeek V4 Pro 0813"]]></title><description><![CDATA[
<p>On DeepSWE it's now 53% vs 63% which is one of the coding benchmarks I trust the most. DS own measurements also show a more significant increase so I suspect AA might update when they release an article.<p>Surprisingly DeepSWE currently shows a lower total cost for pro so that might also update I guess. As usual, don't trust the benchmarks and try for yourself.</p>
]]></description><pubDate>Thu, 13 Aug 2026 17:02:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49288853</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49288853</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49288853</guid></item><item><title><![CDATA[New comment by WASDx in "DeepSeek Harness developer preview"]]></title><description><![CDATA[
<p>I needs to be harness+model combination, <a href="https://artificialanalysis.ai/agents/coding-agents" rel="nofollow">https://artificialanalysis.ai/agents/coding-agents</a></p>
]]></description><pubDate>Thu, 13 Aug 2026 16:54:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49288746</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49288746</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49288746</guid></item><item><title><![CDATA[New comment by WASDx in "ChatGPT Desktop (Codex Desktop) for Linux"]]></title><description><![CDATA[
<p>Right. I'm also curious about those use cases. I don't want an AI clicking through my mailbox.</p>
]]></description><pubDate>Thu, 13 Aug 2026 09:52:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49283678</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49283678</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49283678</guid></item><item><title><![CDATA[New comment by WASDx in "ChatGPT Desktop (Codex Desktop) for Linux"]]></title><description><![CDATA[
<p>I upgraded my workflow a few months ago from "copy-paste things in and out of ChatGPT" to "use an agent that edits my project files and runs tests on its own" and the ergonomics are just so much better and enables automating bigger tasks. I still monitor everything it does and do manual adjustments so I feel ownership of the code.</p>
]]></description><pubDate>Thu, 13 Aug 2026 08:52:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49283289</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49283289</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49283289</guid></item><item><title><![CDATA[New comment by WASDx in "DeepSeek V4 Pro 0813"]]></title><description><![CDATA[
<p>This was a disappointed to me. Why would I use pro over flash now? Is there some area where the difference is significant?</p>
]]></description><pubDate>Thu, 13 Aug 2026 06:56:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49282555</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49282555</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49282555</guid></item><item><title><![CDATA[New comment by WASDx in "Nvidia Nemotron 3.5 Lightning and NeMo Switchyard"]]></title><description><![CDATA[
<p>If you take 10 turns with a model A, it has to read the cache 10 times and write a lot of tokens (the expensive part). Switching to model B is just prefilling the diff + your new message, which is still just one turn. So total number of turns does not increase for a long session even with many switches.<p>I didn't understand this before reading the sibling comments so I'm not sure I got it fully right but I think the total cost becomes like this:<p>* Input tokens: Pay for both models
* Output tokens: Pay for the model that generates
* Cached tokens: Pay per turn, so in total a weighted average over both models?<p>Since output tokens are the most expensive, I can see how this is an overall win for many use cases as benchmarks also show. The hard part is routing correctly.</p>
]]></description><pubDate>Tue, 11 Aug 2026 21:06:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49264488</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49264488</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49264488</guid></item><item><title><![CDATA[New comment by WASDx in "Muse Code and Muse Spark 1.2"]]></title><description><![CDATA[
<p>And they are all TUI's installed via curl | bash.</p>
]]></description><pubDate>Thu, 06 Aug 2026 06:59:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49193427</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49193427</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49193427</guid></item><item><title><![CDATA[New comment by WASDx in "DeepSeek V4 Flash on a Single AMD MI300X"]]></title><description><![CDATA[
<p>At 830tok/s * 1 hour that's almost 3M tokens which is just $0.54 worth of tokens at Deepseeks current output price.</p>
]]></description><pubDate>Tue, 04 Aug 2026 11:17:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49166980</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49166980</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49166980</guid></item><item><title><![CDATA[New comment by WASDx in "The new rules of context engineering for Claude 5 generation models"]]></title><description><![CDATA[
<p>Do you know why they don't just cache the system prompt for everyone? It seems so wasteful not to.</p>
]]></description><pubDate>Sun, 26 Jul 2026 11:04:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49056833</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=49056833</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49056833</guid></item><item><title><![CDATA[New comment by WASDx in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>DeepSWE and FrontierCode are more realistic if you read up on what they actually measure. But the most realistic is to try it yourself. Benchmarks can only vaguely represent typical usage, and how you judge the result. Giving the same real task you have to a few models will make you understand them better than chasing benchmarks.</p>
]]></description><pubDate>Tue, 21 Jul 2026 16:34:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48994637</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48994637</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48994637</guid></item><item><title><![CDATA[New comment by WASDx in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>Great explanation, thanks!</p>
]]></description><pubDate>Sun, 19 Jul 2026 20:08:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48971290</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48971290</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48971290</guid></item><item><title><![CDATA[New comment by WASDx in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>> On top of that, doing research in the open amortizes the cost.<p>Can you elaborate on this? I appreciate the open models but don't see the economics behind just giving them away like now.</p>
]]></description><pubDate>Fri, 17 Jul 2026 21:51:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48952628</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48952628</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48952628</guid></item><item><title><![CDATA[New comment by WASDx in "OpenAI unveils its first custom chip, built by Broadcom"]]></title><description><![CDATA[
<p>I see only these two possibilities:<p>1. If LLMs keep improving, burning models onto silicon becomes obsolete too fast and is not worth doing. Outcome: We keep getting better LLMs.
2. If LLM improvements slow down, they will be burned onto silicon. Outcome: We get faster, cheaper and energy-efficient LLMs.<p>Either way sounds great to me. It will certainly be a mix so we can even get both.</p>
]]></description><pubDate>Wed, 24 Jun 2026 20:31:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48665260</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48665260</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48665260</guid></item><item><title><![CDATA[New comment by WASDx in "GLM-5.2 is the new leading open weights model on Artificial Analysis"]]></title><description><![CDATA[
<p>Are you suggesting it should summarize the image in text or generate it in HTML or something else?</p>
]]></description><pubDate>Wed, 17 Jun 2026 16:19:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48572627</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48572627</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48572627</guid></item><item><title><![CDATA[New comment by WASDx in "Running local models is good now"]]></title><description><![CDATA[
<p>Looking at some benchmarks, the latest ~30B Gemma/Qwen score similar as Claude or GPT versions that were released just <i>one year earlier</i>. That's crazy progress. I can't imagine how it will be in a few years.</p>
]]></description><pubDate>Tue, 16 Jun 2026 17:30:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48558759</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48558759</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48558759</guid></item><item><title><![CDATA[New comment by WASDx in "Open source AI must win"]]></title><description><![CDATA[
<p>I think this is inevitable. Sooner or later, model-specific ASIC's will make economical sense. We're already seeing it happening with Taalas/Cerebras so I think it's sooner than 5 years. And inference is order of magnitude faster which is amazing.</p>
]]></description><pubDate>Sat, 13 Jun 2026 19:45:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48520735</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48520735</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48520735</guid></item><item><title><![CDATA[New comment by WASDx in "Open source AI must win"]]></title><description><![CDATA[
<p>> distributed LLM inference<p>This seems extremely inefficient considering data transfer between model layers if the model is distributed. I found this project called Petals that claim up to 4 tok/s for a 180B model although its repository hasn't been updated in two years.<p><a href="https://petals.dev/" rel="nofollow">https://petals.dev/</a></p>
]]></description><pubDate>Sat, 13 Jun 2026 19:25:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48520543</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48520543</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48520543</guid></item><item><title><![CDATA[New comment by WASDx in "Claude Fable 5"]]></title><description><![CDATA[
<p>I like this one, although its data seem to overlap with ECI.<p><a href="https://artificialanalysis.ai/trends" rel="nofollow">https://artificialanalysis.ai/trends</a></p>
]]></description><pubDate>Tue, 09 Jun 2026 20:15:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=48467049</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48467049</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48467049</guid></item><item><title><![CDATA[New comment by WASDx in "Real-time LLM Inference on Standard GPUs: 3k tokens/s per request"]]></title><description><![CDATA[
<p><a href="https://chatjimmy.ai/" rel="nofollow">https://chatjimmy.ai/</a> from Taalas also feels like that.</p>
]]></description><pubDate>Fri, 29 May 2026 16:44:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48325702</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48325702</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48325702</guid></item><item><title><![CDATA[New comment by WASDx in "Claude Opus 4.8"]]></title><description><![CDATA[
<p>I think their "code" ranking is biased towards visual aesthetics more than raw coding as the voters are just asked which generated website they prefer.</p>
]]></description><pubDate>Thu, 28 May 2026 22:02:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48316137</link><dc:creator>WASDx</dc:creator><comments>https://news.ycombinator.com/item?id=48316137</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48316137</guid></item></channel></rss>