<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: calebkaiser</title><link>https://news.ycombinator.com/user?id=calebkaiser</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 21 Jul 2026 19:26:03 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=calebkaiser" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by calebkaiser in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>I am strongly in favor of open models, open source ML more broadly, and am pretty critical of the cynical positions adopted by major US labs vis a vis open models.<p>But this is an insane characterization. Literally every single researcher and executive at OpenAI and Anthropic would say that "these Chinese labs have a lot of talented people." They hire from them (and vice versa). Tencent's chief AI scientist was poached directly from Deepmind, who poached him from Anthropic, etc etc etc. Do you think there are just zero people from China working at US frontier labs?<p>And even beyond that, the entire ML ecosystem (including people at OpenAI and Anthropic) get excited about research published by Chinese labs. Deepseek's GRPO paper set the ecosystem on fire for a little while.<p>The contention from OpenAI and Anthropic around distillation has basically been "Labs that distill from us get to bootstrap their model at a much lower price point". Or, in other words, "If we didn't invest in building the teacher model, it wouldn't be possible for these labs to distill their student model." Which I'm not very sympathetic to, but is a far cry from how you're characterizing it.</p>
]]></description><pubDate>Mon, 20 Jul 2026 20:37:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48984572</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48984572</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48984572</guid></item><item><title><![CDATA[New comment by calebkaiser in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>OpenAI's Head of Strategic Futures just this week posted this about the latest Kimi release: "It's a very good model! I don't think its performance can be explained away by distillation or anything like that."<p>It was part of a longer post that kicked off quite a firestorm about open models and OpenAI's position on them, but it's also notable that labs are no longer contending that open models are essentially just distilled versions of frontier models: <a href="https://x.com/deanwball/status/2078133895766114412" rel="nofollow">https://x.com/deanwball/status/2078133895766114412</a></p>
]]></description><pubDate>Mon, 20 Jul 2026 20:22:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48984373</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48984373</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48984373</guid></item><item><title><![CDATA[New comment by calebkaiser in "Perfection is not over-engineering"]]></title><description><![CDATA[
<p>This was my initial reaction to reading this post as well.<p>Additionally, as I get older, I find the sentiment of "we're not trying to build a perfect system here" is less about "let's just go fast vroooom" and more akin to saying "I've been humbled before by thinking I had the perfect mental model of the universe before a single user touched the product."</p>
]]></description><pubDate>Mon, 20 Jul 2026 16:27:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48981087</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48981087</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48981087</guid></item><item><title><![CDATA[New comment by calebkaiser in "The state of open source AI"]]></title><description><![CDATA[
<p>This is not my experience working in the field the last 8 years. There is not a dearth of talented researchers, engineers, etc. who are willing to contribute to open models. Just look at the ecosystem generally, from academia to industry. So much research still happens in the open, and the open source community in ML is still massive and energetic.<p>Training a new frontier model just costs a lot of money and there are very few companies for whom it makes sense to do so, and even fewer orgs who are just going to gift those kinds of resources to an open source initiative.</p>
]]></description><pubDate>Fri, 17 Jul 2026 22:10:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=48952770</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48952770</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48952770</guid></item><item><title><![CDATA[New comment by calebkaiser in "Guardian Angels: LLM Personalization for Productivity and Security"]]></title><description><![CDATA[
<p>Gwern's absurdly catalogued personal site is one of those online artifacts that I hope never changes.</p>
]]></description><pubDate>Wed, 15 Jul 2026 02:48:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48915649</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48915649</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48915649</guid></item><item><title><![CDATA[New comment by calebkaiser in "Apple sues OpenAI, accuses ex-employees of stealing trade secrets"]]></title><description><![CDATA[
<p>Based on a cursory read of the situation, it seems similar (at least on its face) to the Waymo vs Uber situation. In that case, Uber payed a Waymo an equity stake and signed an agreement about which technology they would/wouldn't use. The key person involved also was sentenced to 18 months in prison (pardoned after 6 months).</p>
]]></description><pubDate>Sat, 11 Jul 2026 03:52:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=48868565</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48868565</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48868565</guid></item><item><title><![CDATA[New comment by calebkaiser in "GPT-5.6 Sol Ultra produces proof of the Cycle Double Cover Conjecture [pdf]"]]></title><description><![CDATA[
<p>I mean, OpenAI delayed the public release of GPT-2 back in 2019 because it seemed capable of authoring interesting blog posts (that also happened to be untrue). It was a pretty big deal the first time Transformer models were capable of generating that kind of output--no one found it weird. We've just grown to take it for granted that large Transformer models are this capable.<p>The same cycle is happening now for a harder frontier. And proofs represent a pretty good benchmark for model capabilities, so a new model proving a result that a previous model didn't is generally notable in the same way that a model scoring higher on a benchmark is.<p>I'm sure we'll take it for granted in the not-too-distant future.</p>
]]></description><pubDate>Fri, 10 Jul 2026 20:02:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=48864542</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48864542</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48864542</guid></item><item><title><![CDATA[New comment by calebkaiser in "After 7 years in production, Scarf has reluctantly moved away from Haskell"]]></title><description><![CDATA[
<p>I think the author largely agrees with you re: type systems and LLMs. He's pretty explicit that Haskell should be very well positioned to be a power language for LLM-assisted programming, but that the Haskell ecosystem presents the bottlenecks that make it harder.<p>I don't personally use Haskell for anything, but I use Lean and occasionally some other languages with expressive type systems, and like you I've found it to be a pretty great experience for working with LLMs. But I've also experienced what the author is talking about, with languages that sit at different points on the type system spectrum, regarding a languages ecosystem/infra layer becoming a bottleneck. I don't think it's ultimately about the type system but the broader ergonomics of the language/ecosystem.<p>So I think his criticism is less than expressive type systems are a pre-LLM concept, and more that Haskell has an individually bad "agentic coding story".</p>
]]></description><pubDate>Fri, 10 Jul 2026 14:31:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48860528</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48860528</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48860528</guid></item><item><title><![CDATA[New comment by calebkaiser in "After 7 years in production, Scarf has reluctantly moved away from Haskell"]]></title><description><![CDATA[
<p>I've been a power user of LLMs for software development for a while now, and I've found two things to be true:<p>- The benefits of more "extreme" type systems are more accessible and valuable than ever. I have a fairly involved project built on Lean that I hope to open source this month, and it's been a joy to work in even for uses outside of mathematics.<p>- Readability, build time, infra complexity, and everything that affects your speed after finishing your implementation--these things now matter more than ever.<p>It's sort of a dual ergonomics problem, in some sense. And given that, the author's lament makes complete sense to me, especially:<p>"An AI-enabled Haskell ecosystem would ask different questions. How do we make Haskell easier for agents to use well? How do we get more high-quality Haskell examples into model training data? How can we scale reviews? How do we make library docs full of copy-pastable, realistic examples, not just beautiful types? How do we make project bootstrap fast? How do we make error messages more agent-friendly? How do we reduce cold build times? How do we make common industrial patterns obvious to a model that is trying to help?"</p>
]]></description><pubDate>Fri, 10 Jul 2026 14:25:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48860426</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48860426</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48860426</guid></item><item><title><![CDATA[New comment by calebkaiser in "60% Fable cost cut by converting code to images and having the model OCR it"]]></title><description><![CDATA[
<p>Lots of researchers have done just this! There's a really rich history of research + lots of contemporary work on different encoding/representation strategies. This might be interesting to you: <a href="https://sbert.net/" rel="nofollow">https://sbert.net/</a><p>What makes the DeepSeek-OCR and related results exciting to some researchers is less about the fact that you could devise a tokenization scheme that has fewer tokens, and more about how well it works.</p>
]]></description><pubDate>Fri, 03 Jul 2026 21:22:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=48780152</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48780152</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48780152</guid></item><item><title><![CDATA[New comment by calebkaiser in "60% Fable cost cut by converting code to images and having the model OCR it"]]></title><description><![CDATA[
<p>Nah, optical compression is a thing. You see it in a lot of different areas in ML. In this case, the "trick" has been known for a while, and belongs to a whole world of compression research. But I think where you're maybe getting mixed up is in where that 60% gain is coming from.<p>It's not a 60% percent reduction in cost for 100% of the same output. If you have a model and input text A, and you fix the seed etc. and run Text A through the model as text tokens and as compressed image tokens, you will not get identical outputs. You're specifically reducing the number of tensors needed to represent your input, which saves you on raw compute, but also by definition gives you less room to represent the information in your input. It's lossy, in other words.<p>Put another way, if you're using a model like Fable because you need the absolute frontier of capability and cheaper models cannot solve your tasks, then there is a very real chance that a compression strategy like this drops Fable's accuracy such that it's no longer suitable for your task. Which defeats the point of you paying for the most expensive model in the first place.<p>So, it's cool research. Might be useful for some people. Probably isn't something that has incredible utility in real use cases.</p>
]]></description><pubDate>Fri, 03 Jul 2026 18:26:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48778203</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48778203</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48778203</guid></item><item><title><![CDATA[New comment by calebkaiser in "The gap between open weights LLMs and closed source LLMs"]]></title><description><![CDATA[
<p>This has been a (noble) goal of lots of different projects in the community for a long time. Federated learning projects like Flower have been chipping away at it for a long time. There are many many hurdles to be cleared before anything in this area is super feasible as an alternative, but I applaud everyone who works on it.</p>
]]></description><pubDate>Fri, 26 Jun 2026 23:01:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48693109</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48693109</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48693109</guid></item><item><title><![CDATA[New comment by calebkaiser in "Rio de Janeiro's "homegrown" LLM appears to be a merge of an existing model"]]></title><description><![CDATA[
<p>This is a good starting point: <a href="https://huggingface.co/docs/peft/developer_guides/model_merging" rel="nofollow">https://huggingface.co/docs/peft/developer_guides/model_merg...</a><p>But yes, in general, merging refers to techniques that directly blend the weights of different models mathematically. It had a big moment of popularity ~2 years ago, with many so-called "Frankenmodels" popping up on leaderboards.<p>I tend to think of merging as belonging to the same general umbrella as things like "abliteration", or other techniques that surgically modify the weights of a model without a traditional training/tuning loop. Maxime Labonne is a great person to follow if you're interested in this general area.</p>
]]></description><pubDate>Sun, 14 Jun 2026 18:08:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=48530576</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48530576</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48530576</guid></item><item><title><![CDATA[New comment by calebkaiser in "My Agent Skill for Test-Driven Development"]]></title><description><![CDATA[
<p>I don't understand this line of criticism exactly. By putting new information in the context window, you are materially changing the activations at your point of sampling, which is literally "customizing with mere markdown files."<p>Taken to the extreme, the attitude that there is some special incantation that will unlock all capabilities is silly, and a lot of the "prompt engineering" discourse is similarly kind of dumb, but in-context learning is clearly a real thing.</p>
]]></description><pubDate>Fri, 05 Jun 2026 20:35:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48417876</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48417876</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48417876</guid></item><item><title><![CDATA[New comment by calebkaiser in "IBM Spins Off the First Pure-Play Quantum Chip Foundry"]]></title><description><![CDATA[
<p>Eh, Watson was a classic open domain QA system originally, no deep learning or much of what we think of in an "AI platform" today. It was one of a bunch of such systems that were built in that early 2000s period. They all failed because the approach fundamentally didn't work very well.<p>Here's a write up of some relevant history if you're curious <a href="https://liweinlp.com/1465" rel="nofollow">https://liweinlp.com/1465</a></p>
]]></description><pubDate>Mon, 25 May 2026 19:34:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48270731</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48270731</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48270731</guid></item><item><title><![CDATA[New comment by calebkaiser in "The Companies Cutting Headcount for AI Will Lose to the Ones Who Didn't"]]></title><description><![CDATA[
<p>My experience has been that this is not unique to tech, and is common in all large enough industries. I think it's just the natural emergence of reward hacking i.e. if you're an executive at Pepsi and your job is largely to increase the stock price, and you know that you can do something to change the way your numbers are presented such that Wall St will like it, you'll likely do it.<p>I do think tech certainly has its own flavor though, particularly because of how differently it is treated by investors.</p>
]]></description><pubDate>Fri, 22 May 2026 14:13:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48236136</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=48236136</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48236136</guid></item><item><title><![CDATA[New comment by calebkaiser in "China blocks Meta's acquisition of AI startup Manus"]]></title><description><![CDATA[
<p>If I'm remembering right, it was weirder than that, as Llama's originally release strategy was sort of bizarre.<p>You did have to apply for access, but if you met their criteria (basically if you were the right profile of researcher or in government), you got direct access to the model weights, not just an API for a hosted model. So access was restricted, but the full weights were shared.<p>I believe that the model was leaked by multiple people, some of which didn't work at Meta but had been granted access to the weights.</p>
]]></description><pubDate>Mon, 27 Apr 2026 19:58:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=47926540</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=47926540</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47926540</guid></item><item><title><![CDATA[New comment by calebkaiser in "Google plans to invest up to $40B in Anthropic"]]></title><description><![CDATA[
<p>2 years? 2 years ago, gpt-4o was OpenAI's flagship model. The gap is real, but much smaller than 2 years.</p>
]]></description><pubDate>Sat, 25 Apr 2026 03:46:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=47898458</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=47898458</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47898458</guid></item><item><title><![CDATA[New comment by calebkaiser in "GitHub's fake star economy"]]></title><description><![CDATA[
<p>I've worked on two open source infrastructure projects that raised money now, and am friends with people involved in many more. I'd put a couple of asterisks next to the claims in this article:<p>- VCs definitely cared about our Stars, especially in early stages, but not as our primary metric. I suppose Stars might be the primary metric if they're truly off the charts, but usually they're just one of many social proof signals an investor might look at.<p>- Investors, especially at the earliest stages, are quite a varied bunch. Some were diligent about looking at who was leaving Stars on the repo (i.e. are these accounts fake/do they belong to potential future customers). Some less so. This is true for basically every metric (see: startups that grossly misreport ARR)<p>- Fake GitHub stars were a thing way before 2022. I'd have to look in more detail at the methodology here, but I'd question any analysis that finds that paying for GitHub Stars (or any social following kind of metric) is a strictly post-2022 thing. Any metric that can be construed as social proof will immediately have its own grifter economy. Investors know this and (mostly) do their diligence.<p>Finally, showing numbers is hard for an early stage open source startup. At later stages, you should be able to show an actual business with typical metrics, but at the seed stage you often just have a repo and a website. Your goal is just to get a lot of people using your software. You can add telemetry to track that, but that's a thorny decision. GitHub Stars aren't a terrible proxy for popularity, provided that you audit the quality of the following. A project with a lot of organic stars and forks is, at the very least, a project that a lot of people are familiar with.<p>I'm not saying that GitHub Stars aren't wildly overvalued or gamed, but contextualized properly, they're a reasonable metric to consider, particularly at earlier stages. Most investors aren't just throwing millions at random repositories with 20k Stars from obviously spam accounts.</p>
]]></description><pubDate>Mon, 20 Apr 2026 17:48:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=47837993</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=47837993</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47837993</guid></item><item><title><![CDATA[Opik – The missing observability layer for OpenClaw]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/comet-ml/opik-openclaw">https://github.com/comet-ml/opik-openclaw</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47398709">https://news.ycombinator.com/item?id=47398709</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 16 Mar 2026 13:24:04 +0000</pubDate><link>https://github.com/comet-ml/opik-openclaw</link><dc:creator>calebkaiser</dc:creator><comments>https://news.ycombinator.com/item?id=47398709</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47398709</guid></item></channel></rss>