<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: fnordpiglet</title><link>https://news.ycombinator.com/user?id=fnordpiglet</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Mon, 27 Jul 2026 21:42:51 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=fnordpiglet" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by fnordpiglet in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>I suggest downloading these frontier models just to have a copy; even though it’s 1.5TB, it’s worth sticking in a cheap disk and putting aside. Seeding torrents would be even more useful. The man is coming to lock these down, like they tried to do with encryption algorithms. The only way open software survives regulation is through distribution.<p>Over time the enormous investment in techniques and hardware manufacturing will almost certainly make these runnable in a more practical way. It will be a shame if by the time we get there it’s illegal to distribute them and you have to pay a reg capture  premium and feed the machine.</p>
]]></description><pubDate>Mon, 27 Jul 2026 17:55:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49073251</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=49073251</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49073251</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Elevated errors on Claude Opus 5"]]></title><description><![CDATA[
<p>Yes, it’s way off field in many things, gets into weird minutia without seeing a way out, and it’s often seeing a clear sequence of work but then halts on a statement like “ok I’m going to start now.”  Then after expiring the cache when I notice and ask why didn’t it the response is “no reason starting now!”<p>I see this behavior constantly in 5 - the quality of opus and fable have degraded constantly since 4.6 was such a riotous success</p>
]]></description><pubDate>Mon, 27 Jul 2026 17:48:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49073177</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=49073177</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49073177</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>I have to say cc opus 5 is abysmal. It talks to itself incessantly, gets stuck in minutia, fails to understand problems clearly and makes steering mistakes constantly. It also has a weird behavior where it says “ok I know exactly what to do and I will start now,” then sits waiting for user input.  If you’re not on the ball you’re constantly losing 5m/1h cache. Just give me back 4.6.</p>
]]></description><pubDate>Mon, 27 Jul 2026 17:40:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49073056</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=49073056</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49073056</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Wind turbine is being used to produce zero-carbon "green ammonia" fertilizer"]]></title><description><![CDATA[
<p>As well as tons of other things without electricity involved at all!  Don Quixote was never in fear of being electrocuted.<p>In other news, no whales were bothered by this windmill, either.</p>
]]></description><pubDate>Sat, 25 Jul 2026 16:31:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49048994</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=49048994</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49048994</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>When I say distill I also mean mine it architecturally for insights but I doubt seriously the model training is entirely distillation of Claude, it’s almost certainly a mixture of both original corpus and reinforcement as well as distillation. I think it’s a little condescending to imply that these new open models are cheap ripoffs with nothing original to them. These teams and labs are top tier as well, working under unreasonable constraints imposed by the USG. That’s a powerful combination for creativity.</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:55:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49000014</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=49000014</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49000014</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>Because the open model might have been distilled from outputs of the closed model doesn’t mean they are architecturally equivalent. There is almost certainly innovations in architecture present in the open models that the closed labs didn’t think of. It also is almost certainly true that they aren’t completely built out of a distilled corpus, that reinforcement is equivalent, etc. Therefore closed labs will also benefit from being able to inspect in totality the architecture, activations, weights, and be able to train against it at scale in an ensemble of other models and their internal work.<p>The only way there is no benefit would be is if the open models are literal copies of the closed model, which unless there was direct theft, is highly improbable.</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:52:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999989</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48999989</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999989</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>Interestingly the reality of the open source release is they are opening up to full distillation by the closed source model providers at a deeper and more fundamental level. If anything the open sourcing will help Anthropic and open ai ladder up faster. Open source has always been about mutual cooperation towards a goal and has never closed the door to commercial success. All the hand wringing about open weight models putting closed providers at a disadvantage doesn’t get what working in the open actually does for commercial interests - it is like science in the open - it enables and lifts all boats. Likewise commercial success doesn’t close the opportunity for competition or more open source work - it’s the economy of activity and competition that matters overall.  When things stagnate is when closer concerns turtle up and collude on not competing for each others turf.<p>The future is good and better for everyone the more work is in the open and the more work is in the commercial space. It’s good all around.</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:10:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999598</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48999598</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999598</guid></item><item><title><![CDATA[New comment by fnordpiglet in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>Come now. The purposes for which the data is used is relevant. I am much more concerned with my internal corporate IP being actively used against me than passively used to train the model. I would also note you and sign agreements that prohibit the collection of data for use as well, which is also one of the key selling points of bedrock.  In the west you can actually enforce such an agreement in court and win.<p>I wouldn’t lump this into a west vs east thing as well. This is particularly PRC.  I feel comfortable doing business in Japan, Korea, Singapore, Thailand, Malaysia, etc. But it requires some particularly strong willfulness to pretend the PRC isn’t actively and structurally built around economic espionage, and funneling IP through PRC for short term economic gain has been one of the primary factors in their growth over the last 30 years. This just scales it faster.<p>I wouldn’t expect the USG won’t compel AI companies in the US to disclose and retain data as well - however it’s not a simple thing, the companies are hostile to it themselves, courts are often unsympathetic to the government, and the “machine” for converting it into actionable economic advantage is non existent - and there’s a very significant human component in that all links in the chain are culturally uncomfortable with such things. While it happens and it’s possible it’s very difficult, fraught, and does not scale. The PRC is the opposite - the courts, government, and business culture are all aligned in the goals and processes.</p>
]]></description><pubDate>Sun, 19 Jul 2026 17:03:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48969794</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48969794</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48969794</guid></item><item><title><![CDATA[New comment by fnordpiglet in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>I think a better framing is the marginal utility of the models capability growth. At a certain point frontier models will only be needed for frontier problems. The demand for that capability will decrease with time. The hand wringing about not understanding is to my mind anthropomorphic - AI of today lack agency and awareness. Even the constructed stuff Anthropic puts out there in the model docs involve contrived scenarios to elicit “scary” behaviors. It’s unclear that as models become more sophisticated whether they’re better at instruction following or not but it certainly feels that way - even if it’s through better alignment or just an artifact of scaling. However I think the malign actors of humans using powerful models for bad stuff isn’t unreasonable to be concerned about.<p>The marginal utility problem is a real one for AI companies. I think the current generations are already saturating marginal utility for 95% of the population. Almost everyone I know outside of my career has no use for a more powerful model. This is a serious problem for the economics of AI and semiconductor investment. This is a bigger problem than Chinese models. It leads to a demand curve problem - that supply outstrips demand.</p>
]]></description><pubDate>Sun, 19 Jul 2026 08:05:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48965910</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48965910</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48965910</guid></item><item><title><![CDATA[New comment by fnordpiglet in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>I’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will be actively mined for IP, trade secrets, MNPI, etc, or anything of use to the Chinese government or Chinese companies. I don’t say this to crap on the Chinese - but this is the playbook for the last 30 years.<p>That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material.  The economics are astounding.<p>Kimi?  The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.</p>
]]></description><pubDate>Sun, 19 Jul 2026 07:31:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=48965738</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48965738</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48965738</guid></item><item><title><![CDATA[New comment by fnordpiglet in "The state of open source AI"]]></title><description><![CDATA[
<p>Lost me here - I was there - I worked on open sourcing communicator at Netscape and starting Mozilla. We were the one company that tried to own it all, but were out spent and out maneuvered by a more insidious company. This totally misunderstands the situation and the history. Plus the “our CTO” and “CTO letter” at the start was too pretentious. For gods sake; CTOs are just engineers that aren’t worth trusting because they’ve either maneuvered their way to their seat or they are serving the investors on their knees.  Any CTO worth a damn knows they aren’t worth a damn.<p>“””
We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community rose up to make sure it never could. Twenty-five years later, someone is running the same play. We bet on open the first time. Open won. Together, we can do it again.
“””</p>
]]></description><pubDate>Sat, 18 Jul 2026 07:17:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48955917</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48955917</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48955917</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>And this demonstrates this benchmark does not necessitate achieving those goals to achieve a perfect score. You seem to miss the point that almost all of math, physics, computer science is built on constructing an objective then cheating to attain it, and that demonstrates that there is an equivalency. Maybe the benchmark is flawed, or maybe the goals are not strictly necessary to attain.<p>For instance, is solving a math proof by enumerating all permutations exhaustively on a computer cheating?  Does it matter that it is not a proof by construction?  That its not descriptive?  Of course not. The proof of the four color theorem is all that’s necessary and sufficient to prove it. Calculator at an algebra exam?  Who cares. This isn’t an exam, this is the real world. The fact an AI can use a physics harness to perfectly achieve ARC-AGI-3 without attaining those goals demonstrates the power of the technique and that the goals are not necessary for that class of problems. Then find another benchmark that actually demands the goals be necessary and sufficient to achieve the benchmark goals. But don’t denigrate the fact we have technology today that yesterday was a fantasy.</p>
]]></description><pubDate>Fri, 17 Jul 2026 07:17:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48944263</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48944263</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48944263</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>A great deal of mathematics is transforming nonlinear problems into linear ones and solving them with linear techniques. Others are solving non linear problems through stochastic methods. In almost all cases most non trivial math is done by transforming a harder problem into a simpler one.<p>I get what you mean in terms of testing the model itself to see its improvement in some domain. However if you can transform the domain to be better adapted to the model and achieve the desired results, this is indeed an accomplishment because a whole domain of problems is shown to be practically feasible with this technique without expensive model improvements. Of course the benchmark still exists without the harness, but the harness also exists which allows these problems to be solved.<p>As noted elsewhere the models themselves were used to build the harness, which means the models can in fact score this scores without intervention but building a harness for themselves adapted to the domain and using it. Is this cheating by the goal posts you’re setting?<p>There’s a real tension between “I want to solve problems and this technique shows how to solve the problem domain,” and the “I want to measure how something performs unassisted with other techniques.”  Fortunately it’s not a mutually exclusive situation. You can do both simultaneously, gain the benefit of the technique to transform the problem into something tractable and keep measuring using the benchmark.</p>
]]></description><pubDate>Thu, 16 Jul 2026 20:16:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48939652</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48939652</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48939652</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Stripe and Advent have made a joint offer to acquire PayPal – sources"]]></title><description><![CDATA[
<p>2008 in the EU and last year in the US.</p>
]]></description><pubDate>Wed, 15 Jul 2026 23:37:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=48928632</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48928632</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48928632</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Stripe and Advent have made a joint offer to acquire PayPal – sources"]]></title><description><![CDATA[
<p>Not necessarily - the compliance side would open up stripe entirely unless they very carefully structured as a bank holding company. Stripe has been trying to gain a charter in Georgia for years and just got it last year. I think this taught them it would be easier to acquire one. Additionally PayPal has been a EU bank for 12 years and just recently got its US banking charter.  A global bank is pretty compelling.</p>
]]></description><pubDate>Wed, 15 Jul 2026 23:36:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48928629</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48928629</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48928629</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Stripe and Advent have made a joint offer to acquire PayPal – sources"]]></title><description><![CDATA[
<p>What I find interesting about this isn’t acquiring the flow from PayPal but the fact that PayPal has a bank charter and stripe doesn’t (not a real one, all that Georgia bank stuff is shenanigans). Properly structured this gives stripe a lot of interesting opportunities that they’ve had to delegate to partners and that dilutes not just their margins in the fee stack but restricts the types of transactions  and merchants they can host, beyond the minimal regulatory constraint.<p>The next move would be to offer a stripe managed card network that captures the full MDR interchange fee stack, from issuance to processing to network to bank. They could offer deep discounts to merchants hosting their accounts on their bank issuing cards through them and processing through them by cutting out every middle man. It would instantly become the third largest network in terms of points of sale behind visa and Mastercard. They could further carve out heavy rewards for customers through direct relationships with airlines, hotels, and merchants through their existing issuer services. They would become the first fully integrated vertical in payments, and the efficiency that would give them in fees would allow them to crush the hodge podge payment stacks out there that exist. I’d wager in five years they would dominate the market.</p>
]]></description><pubDate>Wed, 15 Jul 2026 22:03:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48927676</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48927676</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48927676</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Claude Code May–July 2026 weekly limits promotion"]]></title><description><![CDATA[
<p>Agreed. It’s also a lot better at following instructions. If anything codex can get caught into an overly literal adherence to instruction while Claude you can barely trust it to sit still for 5 minutes.  Tell Claude to use an MCP for a task, 30% chance it’ll do it. Provide it skills, 10% chance it’ll use them when appropriate. Codex is almost the mirror of that. It’ll almost always use the MCP and recall the skills.<p>The challenge I think for codex is the restriction on context size and the constantly rolling compactions. They are less aggressive or disruptive but it is still annoying you can’t force a 1mm context window on a 1mm context window model.<p>But it’s recall beyond compaction boundaries is much better than Claude code. The impact of compaction is much less noticeable.<p>Performance on task work is between fable and opus, but the marginal utility between that gap is not enough to pay extra.</p>
]]></description><pubDate>Sun, 12 Jul 2026 18:59:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48883561</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48883561</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48883561</guid></item><item><title><![CDATA[New comment by fnordpiglet in "AI boosts research careers but narrow the span of ideas explored: study"]]></title><description><![CDATA[
<p>Yeah that was my instinct too. What sort of career defining trends are visible with this much historical data?  Feels like someone wrote clickbait research to get published.</p>
]]></description><pubDate>Sun, 12 Jul 2026 18:22:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48883236</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48883236</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48883236</guid></item><item><title><![CDATA[New comment by fnordpiglet in "Claude Code May–July 2026 weekly limits promotion"]]></title><description><![CDATA[
<p>Hmm ok. The fact 5.6 Sol performs around Fable level and is included without mega token spend in the subscriptions means I’ve promoted codex to my primary harness and model. The latest release of the CLI, app, and desktop fills a lot of the gaps.<p>Anthropic painted itself into a corner with fable at many turns and this latest twist is one of the more interesting. Either fable is too expensive to run at scale, or they’re trying to incentivize mega spend on tokens, or whatever - but them locking the frontier model away for the few enterprises willing to spend top dollar while codex is including frontier in the subscription (and I’ve found it also is both less token hungry and the limits are much higher for codex) has finally made me put Claude aside and use it as my backup for very specific tasks, where codex has filled that spot for a long time now.<p>50% more weekly limit, but no fable. Ok. I might have a refactoring job somewhere for you Claude for those extra tokens.</p>
]]></description><pubDate>Sun, 12 Jul 2026 18:19:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48883214</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48883214</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48883214</guid></item><item><title><![CDATA[New comment by fnordpiglet in "GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]"]]></title><description><![CDATA[
<p>But a proof isn’t an explanation it’s a proof. Proof by assuming the opposite is true and demonstrating a contradiction is very indirect and not at all directly explanatory yet it’s a proof non the less. The goal of proofs is to demonstrate something to be provably true, not expository knowledge gathering.<p>In fact most mathematicians (myself included!) think the more clever the trick the better the proof!  The trick itself being clever is interesting because it often yields a new way of tackling or thinking about your own proofs. A bland explanatory proof that elicits some conceptually “why” is only preferable if it has a reason for doing so - does understanding why yield a new avenue of research?  Often then the “why” is quite a clever trick too.<p>I think it’s a bit the opposite of programming. There you want your solutions to demonstrably not be clever and the code be its own documentation. It’s a different discipline.</p>
]]></description><pubDate>Sat, 11 Jul 2026 18:10:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=48874253</link><dc:creator>fnordpiglet</dc:creator><comments>https://news.ycombinator.com/item?id=48874253</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48874253</guid></item></channel></rss>