<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mediaman</title><link>https://news.ycombinator.com/user?id=mediaman</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 14 Aug 2026 13:20:54 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mediaman" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mediaman in "Mistral Patent for “Code implemented tool calls”"]]></title><description><![CDATA[
<p>That poster mixed it up with trademarks, for which enforcement is required to maintain its validity.</p>
]]></description><pubDate>Mon, 10 Aug 2026 17:23:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=49246833</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=49246833</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49246833</guid></item><item><title><![CDATA[New comment by mediaman in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>I don't see how this follows. The cost of nails has fallen by 95% over the last century. It's because the cost of manufacturing has fallen. Not because they are selling the information of nail consumers.<p>Tokens are not normal software, because they have marginal cost, and I think people who are used to software economics really struggle with this. With token generation there really can be manufacturing cost efficiencies where one producer is just straight up better at serving product at a lower marginal cost.</p>
]]></description><pubDate>Thu, 30 Jul 2026 18:37:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49113892</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=49113892</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49113892</guid></item><item><title><![CDATA[New comment by mediaman in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>Yes but there is a big, big market for subagents to consume lots of tokens cheaply and condense information up to parent agents. Luna would not be my choice for planning. But an explorer to comb through a codebase to find relevant parts? Or for enterprise retrieval, where it needs to search across many different types of data to see where to focus efforts for a smarter model? Or to wake up periodically to evaluate some conditions and determine if a bigger model should be spun up? Definitely.<p>I've previously found flash (for all the hate it gets) to be good for these kinds of things. Haiku was fine but it's ancient.</p>
]]></description><pubDate>Thu, 30 Jul 2026 18:34:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49113843</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=49113843</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49113843</guid></item><item><title><![CDATA[New comment by mediaman in "Launch HN: Tokenless (YC S26) – Automatic model switching to save money"]]></title><description><![CDATA[
<p>So this only switches models if the cache is cold, because otherwise the economics of switching don't work. But most agentic work involves long strings of successive tool calls that benefit from a hot cache. Hot cache calls reduce input cost by 90%. This can basically only deliver cost savings in turns where the AI delivers a result to the user, the user waits at least 5 minutes (or the length of the cache), and then responds.<p>But user->AI calls are very much the rare case now, the more agentic the workload. Most of them will be tool->result->tool without the user involved. And token burn is highest with these long running agentic chains, but that's precisely where routing doesn't work because of the KV cache.<p>How do you deal with that?</p>
]]></description><pubDate>Wed, 29 Jul 2026 16:08:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49099358</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=49099358</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49099358</guid></item><item><title><![CDATA[New comment by mediaman in "Judge approves $1.5B Anthropic settlement for pirated books used to train Claude"]]></title><description><![CDATA[
<p>None of this matters, this is the judge approving a voluntary settlement reached between the parties last year.<p>If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.</p>
]]></description><pubDate>Wed, 22 Jul 2026 00:36:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49000315</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=49000315</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49000315</guid></item><item><title><![CDATA[New comment by mediaman in "My USB Drive Has a Hidden Encrypted Vault"]]></title><description><![CDATA[
<p>He's saying they just plug your USB stick into some gizmo sold to the state, and it's going to find enough to escalate it.<p>If the idea is that there's no point in encrypting anything on a USB stick, and you just hope they don't look at it at all, sure.<p>But there is no threat model here that includes "check the USB stick" and does not automatically lead to finding and breaking "hidden" content.</p>
]]></description><pubDate>Tue, 21 Jul 2026 20:58:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48998220</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48998220</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48998220</guid></item><item><title><![CDATA[New comment by mediaman in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>Because they don't have a lot of parameters to store general Wikipedia knowledge. They're small. Use big models that have high parameter capacity to store general information. Or build a harness around the small model that searches a knowledge base/internet.<p>Use the right tool for the job. It's like asking why a screwdriver isn't good at sawing wood, or calling C a terrible language because it's hard to make CRUD apps with it.</p>
]]></description><pubDate>Tue, 21 Jul 2026 19:17:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48996846</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48996846</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48996846</guid></item><item><title><![CDATA[New comment by mediaman in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>Small open source models shouldn't be used for world knowledge, that's not their purpose.</p>
]]></description><pubDate>Tue, 21 Jul 2026 17:53:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=48995727</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48995727</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48995727</guid></item><item><title><![CDATA[New comment by mediaman in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>Gemma 4 was released in April. It's a good series of multimodal models.</p>
]]></description><pubDate>Tue, 21 Jul 2026 16:44:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48994797</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48994797</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48994797</guid></item><item><title><![CDATA[New comment by mediaman in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>There was some recent reporting that a July release of the Pro model got pushed back for exactly that reason. Its performance was not good compared to the OpenAI/Anthropic big models. They are having a lot of problems with posttrain.</p>
]]></description><pubDate>Tue, 21 Jul 2026 16:35:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48994649</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48994649</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48994649</guid></item><item><title><![CDATA[New comment by mediaman in "Who's afraid of Chinese models?"]]></title><description><![CDATA[
<p>You're correct, but that's a different market segment and not the market GLM 5.2 and its peers compete in.<p>The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.<p>From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.<p>But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.<p>That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)</p>
]]></description><pubDate>Tue, 21 Jul 2026 02:03:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48987216</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48987216</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48987216</guid></item><item><title><![CDATA[New comment by mediaman in "Who's afraid of Chinese models?"]]></title><description><![CDATA[
<p>The (quite excellent) article discusses several of your points. If you haven't read it, I recommend it.<p><pre><code>  - Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.

  - Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use

  - The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6.

  - This is because US labs are leading on cost efficacy of inference ($/task)

  - Training will decline as a percentage of costs as inference expands compute share due to agentic workloads. A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.

  - With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
</code></pre>
OpenAI really shows the way here. Their cost per task is less than half that of Anthropic because of more efficient tokenization and less verbosity. OpenAI is both cheaper and better than Chinese models for frontier work.</p>
]]></description><pubDate>Mon, 20 Jul 2026 23:04:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48985989</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48985989</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48985989</guid></item><item><title><![CDATA[New comment by mediaman in "Who's afraid of Chinese models?"]]></title><description><![CDATA[
<p>Convergence in coding makes them highly substitutable. But I could see harnesses configured for different purposes -- let's say, a harness for creating teaching plans -- being able to cater to its audience better than a coding harness. Maybe it's got tools to plug into standardized curricula, what the lesson books will be, what other lesson plans the district's teachers have made, etc., which could be done in a clunky way in a regular harness but could be streamlined.</p>
]]></description><pubDate>Mon, 20 Jul 2026 22:45:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=48985822</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48985822</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48985822</guid></item><item><title><![CDATA[New comment by mediaman in "Who's afraid of Chinese models?"]]></title><description><![CDATA[
<p>This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.</p>
]]></description><pubDate>Mon, 20 Jul 2026 22:41:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48985788</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48985788</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48985788</guid></item><item><title><![CDATA[New comment by mediaman in "Billion Dollar PDFs"]]></title><description><![CDATA[
<p>The problem with many of these kinds of works is that they are a product of a provincial SV mindset that is largely unaware of the broader world or anything at all outside of its cultural and economic ecosystem, so it winds up being a super narrow slice of conventional, mostly boring topics that everyone in SV rehashes over and over.<p>Wouldn't it be interesting to learn an influential work that changed how health care professionals run hospitals? Or a document that changed how mining works? A paper on Wright's theory of manufacturing scaling that explains the solar revolution? A thesis on how the world's factories moved from pneumatics to servo systems and why? How a policy thesis changed federal regulators' approach to approving rare disease drugs? Maybe Hayek's views on socialism and information theory, or perhaps an influential thesis on how antitrust monopoly regulation should work?<p>But instead it's a bunch of crypto and AI stuff that everyone in tech already knows about, rewarmed again for its hundredth serving.</p>
]]></description><pubDate>Mon, 13 Jul 2026 01:31:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48886783</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48886783</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48886783</guid></item><item><title><![CDATA[New comment by mediaman in "Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper"]]></title><description><![CDATA[
<p>Because quality of writing matters.<p>Good communicators learn to use the written word. Bad ones rely on mental crutches.<p>Good communicators get an audience, and bad ones won't.<p>You think it's a lost cause, but it's not, because people don't like this junk, because it is low quality and, on average, lacks substance.<p>The best minds in AI that I've seen all write their own words. They use AI to help them research or ideate, but what they write is their own.<p>Before assuming this is a "lost cause," consider why the smartest people in the room don't do it.</p>
]]></description><pubDate>Mon, 13 Jul 2026 01:01:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48886604</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48886604</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48886604</guid></item><item><title><![CDATA[New comment by mediaman in "45% of Enthusiasts 'Seriously Considering' Leaving Sony for PC"]]></title><description><![CDATA[
<p>That's annoying, but it's not why scaled manufacturing is lowering unit costs of panel production. Look at bare panel prices, they've followed the same cost curve down.<p>The same problem exists in the airline market. Airline ticket prices are historically very low, but people complain about seats, fees, and so on. But then they keep buying the absolute cheapest ticket.<p>What consumers say they care about, and what they actually care about, are not the same. Otherwise they'd pay more for the less irritating product.</p>
]]></description><pubDate>Fri, 10 Jul 2026 20:22:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48864752</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48864752</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48864752</guid></item><item><title><![CDATA[New comment by mediaman in "GLM 5.2 is nearly as accurate as a human book keeper"]]></title><description><![CDATA[
<p>I have seen the SWIFT thing happen for $100k. I think AI could actually be better for this, because it's often easier to implement hard rules for the AI.<p>With the SWIFT incident I saw, there was a rule that no payment can go to a vendor's bank that isn't a current, approved vendor. But the rule was not enforced in software: it was an internal accounting rule that humans were supposed to follow. The AP person "thought it had been approved" because there was a similar transaction with a different company that was a new vendor at a similar time. The other transaction was legitimate, the fraudulent spoofer wasn't. The wire got sent to a party in China.<p>With AI agents, if you approach it from the perspective that it will be gullible and trickable by fraudsters, you build in these hard guardrails. With humans, it's much easier to believe that "we trained Lucy on this procedure" will work in all circumstances, even if Lucy still has the technical ability to bypass the official procedure.<p>In these cases, it starts looking a lot more like traditional software, with your little AI chaos monkeys constrained in little boxes within the software chain.</p>
]]></description><pubDate>Thu, 09 Jul 2026 20:55:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48852248</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48852248</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48852248</guid></item><item><title><![CDATA[New comment by mediaman in "GLM 5.2 is nearly as accurate as a human book keeper"]]></title><description><![CDATA[
<p>You can fix this simply by using normal controls.<p>That's why we have purchase orders that can only be entered by buyers. Product is received and approved by buyer. Invoice goes to accounting, who can't approve it unless there's a matching purchase order and receiver.<p>Yes, letting agents do whatever they want leads to disaster. But humans are gullible stochastic token generators as well. And that's why the problem is already solved.</p>
]]></description><pubDate>Thu, 09 Jul 2026 20:43:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48852107</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48852107</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48852107</guid></item><item><title><![CDATA[New comment by mediaman in "Chatto is now open source"]]></title><description><![CDATA[
<p>Looks like it's planned.<p><a href="https://github.com/orgs/chattocorp/projects/1?pane=issue&itemId=202121285&issue=chattocorp%7Cchatto%7C995" rel="nofollow">https://github.com/orgs/chattocorp/projects/1?pane=issue&ite...</a></p>
]]></description><pubDate>Wed, 08 Jul 2026 17:17:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=48834569</link><dc:creator>mediaman</dc:creator><comments>https://news.ycombinator.com/item?id=48834569</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48834569</guid></item></channel></rss>