<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Implicated</title><link>https://news.ycombinator.com/user?id=Implicated</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 08 Oct 2026 05:08:34 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Implicated" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by Implicated in "Claude Code’s suggested message feature: I think the real customer is the model"]]></title><description><![CDATA[
<p>seems like maybe we're all just not as mature as you, unfortunate for you.</p>
]]></description><pubDate>Wed, 07 Oct 2026 04:51:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49988343</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49988343</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49988343</guid></item><item><title><![CDATA[New comment by Implicated in "Mistral Large 4"]]></title><description><![CDATA[
<p>> not to clone someone else's work and release it for less money.<p>... You're saying this about _deepseek_?</p>
]]></description><pubDate>Tue, 06 Oct 2026 16:44:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49980980</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49980980</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49980980</guid></item><item><title><![CDATA[New comment by Implicated in "Turn off Apple Intelligence on macOS 27 and get its disk space back"]]></title><description><![CDATA[
<p>When I read things like this all I can imagine is a farmer responding in a similar fashion to people talking about using tractors.</p>
]]></description><pubDate>Mon, 05 Oct 2026 17:34:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49967857</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49967857</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49967857</guid></item><item><title><![CDATA[New comment by Implicated in "How did AMD Ryzen get 50% faster in two years?"]]></title><description><![CDATA[
<p>I picked up a pre-ai-price-insanity AX162-R at Hetzner a while back and loaded it up on memory to max out the 12 channels the 48c EPYC 9454P as a "this will be the last mysql box I'll need" and have I been _wildly_ impressed with it's performance. The things I throw at it are honestly laughable at times, wildly irresponsible queries against a rather large database, the redis qps metrics are ridiculous and I just load it up with random ggufs since the memory bandwidth is... not terrible and 384GB of it is... useful.<p>The web app that's hosted on it deals with lots of images and text - over 100mil of each deduplicated, embedded, simhashed - it does the hashing, and the embedding in real time during ingest. Just handles it.<p>These things are absolutely insane.</p>
]]></description><pubDate>Wed, 23 Sep 2026 01:26:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49810535</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49810535</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49810535</guid></item><item><title><![CDATA[New comment by Implicated in "MiMo v2.6"]]></title><description><![CDATA[
<p>> I did not personally test the open weight models beyond the old Qwen 3.6 27B, which produced unusably bad results for me.<p>So you don't have much perspective on things, it seems. Let me introduce you to the GLM 5.2 and then 5.3/5.3 flash series of... "oh, wow, I should have bought some  RTX PRO 6000's while they were 'cheap'" stage of progression.<p>As someone carrying multiple max subscriptions to both claude and codex - primary workhorse is glm 5.3 flash running on rented GPUs for less than a latte/hr.<p>I also found qwen 3.6 27B nearly useless for my own needs. DS4 flash 0731 and then 4.1 have been nearly as eye opening as glm 5.3 flash, but have their own warts.</p>
]]></description><pubDate>Mon, 21 Sep 2026 22:02:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49794048</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49794048</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49794048</guid></item><item><title><![CDATA[New comment by Implicated in "Astra for Law"]]></title><description><![CDATA[
<p>> it seems like for law specifically all relevant facts will be cited and checked easily by humans.<p>Don't be too sure about that. [0]<p>0: <a href="https://www.damiencharlotin.com/hallucinations/" rel="nofollow">https://www.damiencharlotin.com/hallucinations/</a></p>
]]></description><pubDate>Thu, 17 Sep 2026 21:04:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49746534</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49746534</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49746534</guid></item><item><title><![CDATA[New comment by Implicated in "Nitter and XCancel resume service after legal advice"]]></title><description><![CDATA[
<p>I don't have a horse in this race but I would question the idea that a municipal service has a better reliability record than x/twitter.</p>
]]></description><pubDate>Mon, 07 Sep 2026 03:22:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49593517</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49593517</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49593517</guid></item><item><title><![CDATA[New comment by Implicated in "GPT-6 Astra"]]></title><description><![CDATA[
<p>> Sol is so much better than Fable 5.<p>... <i>looks around</i> ...</p>
]]></description><pubDate>Thu, 03 Sep 2026 21:58:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49557709</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49557709</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49557709</guid></item><item><title><![CDATA[New comment by Implicated in "Hy4 preview"]]></title><description><![CDATA[
<p>I think you're arguing the same general point that the person you're responding to is. But you're saying he's not understanding - he understands that they report a cache hit % but you can't look at that public metric with any level of accuracy _because_ most people aren't pinning their providers and they _are_ getting juggled around which is bringing that metric down. That's not to say that specific providers might have issues or worse cache implementations - but it stands that if openrouter is juggling the requests back and forth by default then _that alone_ is breaking caches on those requests in huge numbers.</p>
]]></description><pubDate>Sat, 29 Aug 2026 23:52:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49494352</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49494352</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49494352</guid></item><item><title><![CDATA[New comment by Implicated in "Hy4 preview"]]></title><description><![CDATA[
<p>> OpenRouter randomizes which provider gets your request by default right?<p>I'm not sure it's wholey accurate to say they "randomize" the provider, rather my assumption based on usage is that it's something like cheapest-ish/responded to the request within some reasonable-ish time/etc algorithm that chooses the provider on each request - which seems, remarkably questionable in terms of optimizing for user experience or hidden user costs.<p>> This behavior makes it so you don't benefit much from the caching, unless you pin it to a single provider.<p>I so very much recommend this approach. My avenues that automate llm calls to openrouter are setup to make api reqs to openrouter to determine best price/response/etc and then pin the request to that (and, preferably, a fallback if there's reasonable difference between #1 and #2) provider for that session. Otherwise you're going to have a bad time.<p>I'd imagine this could make things interesting in cases where one provider is offering different quants than the others and openrouter is just swapping you back and forth on a long agentic session.</p>
]]></description><pubDate>Sat, 29 Aug 2026 23:49:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49494331</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49494331</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49494331</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>> and it was running very slowly<p>... I'm at a loss for words here. It was being served for free. To the entire world.</p>
]]></description><pubDate>Wed, 26 Aug 2026 17:19:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452662</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452662</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452662</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>> All the American companies you mentioned still follow American law and regulation. Skirting that blatantly has big consequences.
> Chinese companies do not follow American laws and there are absolutely no consequences for violating it.<p>... lmk when anthropic/openai/spacex/xai are held accountable for anything. Anything at all. Hard to be when you're _writing_ the rules.</p>
]]></description><pubDate>Wed, 26 Aug 2026 17:16:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452594</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452594</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452594</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>At the current point in time I'd argue it's more about opportunity cost/value.<p>If I'm a professional photographer chasing the best possible end product, I'm not buying cameras because they're economical. I'm buying the best camera I can get my hands on to get the best product I can produce within reason under the understanding that it doesn't have to equate to the best economic decision to be the _right_ decision.<p>If you're in a position to be able to take advantage of the local inference - it's a no brainer. If you're not sure how that would be done, then it's not a good move.</p>
]]></description><pubDate>Wed, 26 Aug 2026 17:08:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452475</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452475</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452475</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>If your usage wouldn't change with local inference and you don't have security/privacy concerns then at the currently heavily subsidized pricing, sure.. not economical.<p>But things change real fast when you're no longer bound by costs/apis/rate limits. All of a sudden it's not about "how can I do this right and efficiently" and more about "I can poke at and test _all the things_ that might make this better".<p>I think most people who can't see this value in the local inference approach are likely still copy/pasting from their web LLM ui's or don't even come close to subscription quotas. Meanwhile, 1b tokens a day is a light day for me with 3 $200/m subscriptions + some level of sub at basically every frontier level provider. Had I been less frugal and ponied up for the hardware before things got crazy I wouldn't need 80% of that - just the frontier models for the most complex tasks, the open weight models would handle the rest easily _and_ I'd get to do a lot more exploratory work without concern about quotas.</p>
]]></description><pubDate>Wed, 26 Aug 2026 17:06:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452445</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452445</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452445</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>Use Opus 4.8. 5 is absolute garbage.<p>Don't use DS4 Flash in max effort mode. It's just spinning its wheels, in my experience (I have a harness for testing models with 25 real bugs/features/etc from my real projects that I measure outcomes against) DS4 flash does _worse_ with max effort. It will literally have the right approach and reason itself away from it.</p>
]]></description><pubDate>Wed, 26 Aug 2026 17:01:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452374</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452374</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452374</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.</p>
]]></description><pubDate>Wed, 26 Aug 2026 16:57:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452310</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452310</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452310</guid></item><item><title><![CDATA[New comment by Implicated in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>As a counter to that - I've tried various flavors/quants/full weights and Qwen 3.8 27B has been entirely <i>useless</i> at anything non-trivial. Sure - it can do some boilerplate work (though, even armed with a well written spec and working within a very well known framework it went off the rails and did things in a way that were... um... questionable at best) but I don't see it as anything more than a personal assistant style model. Zero chance I'd "work" with it, I spent days trying to get it to do something for me that was usable that I didn't have to have reviewed and refined by a frontier level model or myself. Couldn't do it. The idea that qwen 3.8 27b is _anywhere near_ Opus 4.8 is laughable. Pure benchmaxxing.<p>DS4 Flash 0731, on the other hand, wildly opposite experience. Would recommend.<p>GLM 5.2 - even quanted down to a hybrid 4/3 bit setup is amazing for everything but the hardest/most complex stuff in the same projects/realm.</p>
]]></description><pubDate>Wed, 26 Aug 2026 16:53:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49452262</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49452262</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49452262</guid></item><item><title><![CDATA[New comment by Implicated in "Firefox 157 will include JPEG XL by default on all platforms"]]></title><description><![CDATA[
<p>> JpegXL will be the best of all worlds choice in a few years when software support is good.<p>Where is the software support lacking other than the browsers at this point?</p>
]]></description><pubDate>Tue, 25 Aug 2026 21:25:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49440876</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49440876</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49440876</guid></item><item><title><![CDATA[New comment by Implicated in "Clean up Claude 5's token vomit with a separate LLM"]]></title><description><![CDATA[
<p>Whether or not they can be trusted isn't all that relevant when it's still something along the lines of "Insert $1 get $25 in return" even if it's their own rates you're using to measure the value. I'm at ~2.2b Fable 5 tokens in the last 7 days (I ingest/index every session) and napkin math says that's ~$2,700 in usage. I have two max accounts, so $400 a month, divide by 4 to get $100 for this same 7 day period across those two accounts (neither are maxed out for the week, so this isn't even full utilization). I put $100 into the machine and got back $2,700 in fable bucks. Deepinfra would have to have quite the discounted rate to beat that.</p>
]]></description><pubDate>Thu, 20 Aug 2026 16:55:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49377179</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49377179</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49377179</guid></item><item><title><![CDATA[New comment by Implicated in "Vomit: Clean up Claude 5's token output with a separate LLM"]]></title><description><![CDATA[
<p>> Why not just use that other vendor's model for everything?<p>Because it's not an either or thing. Neither is sufficient. I'd argue that, expenses aside, you should have every model you have access to cross reviewing the work of the others.<p>Outside of super trivial things that I should have just done myself, I have a cross-model review of _everything_ these days. The tokens are too cheap not to.</p>
]]></description><pubDate>Thu, 20 Aug 2026 16:42:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49377008</link><dc:creator>Implicated</dc:creator><comments>https://news.ycombinator.com/item?id=49377008</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49377008</guid></item></channel></rss>