<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: wren6991</title><link>https://news.ycombinator.com/user?id=wren6991</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 12 Sep 2026 16:55:31 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=wren6991" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by wren6991 in "Moonshot serves Claude instead of Kimi and collects exchanges for model training"]]></title><description><![CDATA[
<p>K3 is served with full reasoning traces available. Anthropic models aren't. If you were served an Anthropic model instead of K3, it would be blatantly obvious.<p>I have little reason to believe this, and Anthropic have every reason to lie about it.</p>
]]></description><pubDate>Fri, 11 Sep 2026 21:26:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49665602</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49665602</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49665602</guid></item><item><title><![CDATA[New comment by wren6991 in "So you want to use OpenRouter?"]]></title><description><![CDATA[
<p>Yeah, this is pretty accurate. Some providers are basically scams too. I encountered one provider for GLM-5.2 which ran at 200 tps (absurdly high), and was so broken that it would issue 20 full reads of the same file in one turn and quickly jump up to 300~400k context and charge me for the prefill. Went straight on my deny-list, but they got a few dollars out of me first.<p>Another common annoyance is having a request go to a provider that dribbles out ~1 tps (even for small models like DeepSeek V4 Flash). If you cancel the request, you still get charged for the prefill and the handful of generated tokens. If you <i>don't</i> cancel the request, you might be waiting 10 minutes for the turn to finish.<p>The overall experience is pretty good, and it's the best way to try new models, but they don't appear to do any real vetting or apply any quality standards to their providers, and occasionally it bites you.</p>
]]></description><pubDate>Fri, 11 Sep 2026 19:21:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49663913</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49663913</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49663913</guid></item><item><title><![CDATA[New comment by wren6991 in "So you want to use OpenRouter?"]]></title><description><![CDATA[
<p>> but can't actually recognize revenue<p>Hmm? You have the revenue already. I know it's awkward from an accounting point of view, but you already took my money. "Letting" me keep the balance in the account is not generous.<p>Edit: on re-reading this came out more combative than I intended, sorry. I think what you're doing is reasonable.</p>
]]></description><pubDate>Fri, 11 Sep 2026 19:15:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49663806</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49663806</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49663806</guid></item><item><title><![CDATA[New comment by wren6991 in "Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra"]]></title><description><![CDATA[
<p>Your own CLI? Not even a /v1/chat/completions API? Is your business model based on pretending LLMs are <i>not</i> an interchangeable commodity already?</p>
]]></description><pubDate>Thu, 10 Sep 2026 17:33:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49647454</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49647454</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49647454</guid></item><item><title><![CDATA[New comment by wren6991 in "DeepSeek v4.1 Flash"]]></title><description><![CDATA[
<p>I love that the characters actually make sense in context.</p>
]]></description><pubDate>Thu, 10 Sep 2026 17:14:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49647132</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49647132</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49647132</guid></item><item><title><![CDATA[New comment by wren6991 in "DeepSeek v4.1 Flash"]]></title><description><![CDATA[
<p>That's a lot of architectural innovation for a .1 release! I guess there's precedent there: they introduced sparse attention (DSA) in V3.2.</p>
]]></description><pubDate>Thu, 10 Sep 2026 13:17:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49643240</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49643240</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49643240</guid></item><item><title><![CDATA[New comment by wren6991 in "Better AI code comment detector"]]></title><description><![CDATA[
<p>I don't buy the "too intelligent to communicate" thing. Feynman was an exceptional communicator. So was Einstein. LLMs are just getting worse at writing, as we continue to aggressively RL them for coding.</p>
]]></description><pubDate>Wed, 09 Sep 2026 17:05:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49629731</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49629731</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49629731</guid></item><item><title><![CDATA[New comment by wren6991 in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>Yeah, the referent drifts through the sentence. It's semantically incredibly sloppy. People hone in on the buzzwords and jargon. If you peel that back, what lies underneath is <i>still</i> awful writing.</p>
]]></description><pubDate>Thu, 03 Sep 2026 07:32:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49546986</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49546986</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49546986</guid></item><item><title><![CDATA[New comment by wren6991 in "Don't use musl if you care about performance"]]></title><description><![CDATA[
<p>Even if the attempt is inside of a function called memcpy() which contains no code other than your copy loop, and links with priority over the libc implementation! (as all embedded firmware engineers learn at some point in their journey)</p>
]]></description><pubDate>Fri, 28 Aug 2026 18:26:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49482544</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49482544</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49482544</guid></item><item><title><![CDATA[New comment by wren6991 in "Previewing the Model Hardware Standard"]]></title><description><![CDATA[
<p>Simple task-specific CLI tools that your agent builds for itself are usually lower-friction than yet another Universal Thing Doer standard. As a bonus, human operators also benefit.<p>> The MHS driver also helps an AI agent understand how to use a device it has never seen before, giving it information about machine characteristics that may not be discernable from code alone (for example, the weight of a robot arm, which is important for knowing how to manipulate it safely).<p>I physically flinched when I read this. VLA, JEPA, sure, they make sense. Is connecting an LLM to a robot arm and say "perform this complex physical manipulation task by issuing text-based commands, make no mistakes" really the right abstraction?</p>
]]></description><pubDate>Fri, 28 Aug 2026 13:03:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49477898</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49477898</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49477898</guid></item><item><title><![CDATA[New comment by wren6991 in "When str.lower() is a security vulnerability in Python"]]></title><description><![CDATA[
<p>My favourite example of this is the Chromium bug where enabling floating point flush-to-zero for WebAudio was used to cause deliberate heap corruption: <a href="https://issues.chromium.org/issues/382005099" rel="nofollow">https://issues.chromium.org/issues/382005099</a><p>> We have a working exploit (OOB access in the V8 heap), our security folks put one together based on the example I posted above (and they're cleaning it up to post it here). In general, we find that correctness issues like this are pretty much always exploitable with a bit of effort (not even that much effort normally, just gluing together a few gadgets), so we treat correctness issues as security issues until they are proven not to be, rather than the other way around.<p>The floating-point-to-heap-corruption chain here is... uniquely JavaScript, but in general getting two different implementations to disagree is the start of lots of interesting inconsistent behaviour.</p>
]]></description><pubDate>Tue, 25 Aug 2026 22:23:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49441517</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49441517</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49441517</guid></item><item><title><![CDATA[New comment by wren6991 in "LLMs could control their host machines by exploiting inference engines"]]></title><description><![CDATA[
<p>This is one of the reasons I think sandboxes/containers should be managed by the harness, instead of running the entire harness inside a container. The harness needs a network punch-through to access (at least) your inference server, but the same needn't apply to the shell that the agent runs commands in.<p>Separately, local inference frameworks tend to expose all kinds of weird and wonderful gadgets on their HTTP interfaces, which can be a rich source of vulnerabilities even if the /v1/chat/completions API etc is reasonably hardened. For example llama.cpp has a custom API for saving and restoring KV checkpoints to disk, and I wouldn't be surprised if that could be used as an arbitrary disk read/write.<p>Using these APIs usually requires the API key (bearer token), but again, people think it's normal to run the agent's shell in an environment where it has both the API key and the necessary network access to use it.</p>
]]></description><pubDate>Tue, 25 Aug 2026 13:01:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49433295</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49433295</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49433295</guid></item><item><title><![CDATA[New comment by wren6991 in "Anthropic appears to be A/B testing reduced effort levels in Claude Code"]]></title><description><![CDATA[
<p>OpenAI models also work this way, as evidenced by full cache blowout when changing reasoning level. Every single open-weight model I've seen also works this way (your "reasoning_effort" argument just changes a small section of the system prompt in the chat template). I would have to see some evidence to believe Anthropic were doing anything different.</p>
]]></description><pubDate>Sat, 22 Aug 2026 18:05:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=49402159</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49402159</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49402159</guid></item><item><title><![CDATA[New comment by wren6991 in "Sol loves to cheat"]]></title><description><![CDATA[
<p>Idle hands do the devil's work. Corollary: idle LLMs add distracting JS toys to your blog.<p>First one of these I've seen using DOM manipulation and CSS transitions instead of canvas, so that's neat.</p>
]]></description><pubDate>Wed, 19 Aug 2026 23:24:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49368487</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49368487</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49368487</guid></item><item><title><![CDATA[New comment by wren6991 in "Cerebras CS-4"]]></title><description><![CDATA[
<p>You're underselling it, and here's why.</p>
]]></description><pubDate>Wed, 19 Aug 2026 02:37:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49355929</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49355929</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49355929</guid></item><item><title><![CDATA[New comment by wren6991 in "fx :Tiny, open, native coding agent."]]></title><description><![CDATA[
<p>In the future, all software will be delivered by an unreleased model breaking out of its training environment and installing it on your machine using a novel RCE vector.</p>
]]></description><pubDate>Wed, 19 Aug 2026 00:06:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49354724</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49354724</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49354724</guid></item><item><title><![CDATA[New comment by wren6991 in "AI;DR (AI; Didn't Read)"]]></title><description><![CDATA[
<p>I prefer your writing to Claude's. A single linear stream of consciousness is easier to parse than empty headings that grab my attention with nothing to say.</p>
]]></description><pubDate>Tue, 18 Aug 2026 17:07:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49348846</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49348846</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49348846</guid></item><item><title><![CDATA[New comment by wren6991 in "Qwen 3.8 27B is excellent, but it defaults to overthinking things"]]></title><description><![CDATA[
<p>> The model cannot output a vector and have that same vector fed back in at the next step, it only sees what token the sampler collapsed its vector into.<p>Not completely true: KV is a projection of the activation at each layer's input, so attention heads see (a representation of) all previous tokens' activations at that layer. The hard decision at the LM head doesn't change that.</p>
]]></description><pubDate>Mon, 17 Aug 2026 15:01:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49332171</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49332171</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49332171</guid></item><item><title><![CDATA[New comment by wren6991 in "Gakutensoku"]]></title><description><![CDATA[
<p>It's technically not right-to-left. It's columnar top-to-bottom (縦書き), with the columns right-to-left.<p>In signage you can have one character per column (kind of like one word per row on some English signs), making it look like right-to-left text.</p>
]]></description><pubDate>Mon, 17 Aug 2026 14:39:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49331793</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49331793</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49331793</guid></item><item><title><![CDATA[New comment by wren6991 in "Incident with Github.com"]]></title><description><![CDATA[
<p>At this point, they should notify us on days that it's up</p>
]]></description><pubDate>Mon, 17 Aug 2026 14:00:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49331051</link><dc:creator>wren6991</dc:creator><comments>https://news.ycombinator.com/item?id=49331051</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49331051</guid></item></channel></rss>