<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: croemer</title><link>https://news.ycombinator.com/user?id=croemer</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 11 Oct 2026 07:40:01 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=croemer" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by croemer in "RIP, vector database"]]></title><description><![CDATA[
<p>I'm not sure I've ever read a sentence like this before LLMs. Maybe I didn't pay attention, possible. I should do a search of pre-LLM writing for this construction.</p>
]]></description><pubDate>Sat, 03 Oct 2026 13:03:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49943919</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49943919</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49943919</guid></item><item><title><![CDATA[New comment by croemer in "RIP, vector database"]]></title><description><![CDATA[
<p>I see, you're updating the graph with 3 weeks delay, that's a bit confusing. Now there are 4 days at least.</p>
]]></description><pubDate>Sat, 03 Oct 2026 13:02:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49943907</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49943907</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49943907</guid></item><item><title><![CDATA[New comment by croemer in "Pi Durable"]]></title><description><![CDATA[
<p>That code font is painful to read, no syntax highlighting and extremely pixelated. It's retro but an eyesore.</p>
]]></description><pubDate>Thu, 01 Oct 2026 22:44:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49927929</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49927929</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49927929</guid></item><item><title><![CDATA[New comment by croemer in "RIP, vector database"]]></title><description><![CDATA[
<p>Dashboard was last updated on September 7, it started on September 5. Bug? Or no progress? It's linked in the blog post so would expect it to work: <a href="https://turbopuffer.com/v3" rel="nofollow">https://turbopuffer.com/v3</a></p>
]]></description><pubDate>Thu, 01 Oct 2026 22:40:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49927906</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49927906</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49927906</guid></item><item><title><![CDATA[New comment by croemer in "RIP, vector database"]]></title><description><![CDATA[
<p>This one sentence sounds very LLMish:<p>> Object storage as the source of truth gave the economics, and tiered NVMe SSD/memory caches gave the performance.</p>
]]></description><pubDate>Thu, 01 Oct 2026 22:35:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49927875</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49927875</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49927875</guid></item><item><title><![CDATA[New comment by croemer in "Clef: Open-source decision models, and new RL fine-tuning platform"]]></title><description><![CDATA[
<p>Thanks for the tip! Set it up locally and it works!</p>
]]></description><pubDate>Thu, 01 Oct 2026 18:40:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49925436</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49925436</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49925436</guid></item><item><title><![CDATA[New comment by croemer in "Clef: Open-weight decision models, and new RL fine-tuning platform"]]></title><description><![CDATA[
<p>Is there a Jev-like model I can run on my Mac? Something like Ollama? Or what's the best way to play with it? Is there a cheap/free API service eg on OpenRouter?</p>
]]></description><pubDate>Thu, 01 Oct 2026 18:11:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49925083</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49925083</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49925083</guid></item><item><title><![CDATA[New comment by croemer in "Backblaze drive stats for Q2 2026"]]></title><description><![CDATA[
<p>Oh cool, I missed that the raw data is available. I shall see if I can do the analysis myself then.</p>
]]></description><pubDate>Thu, 01 Oct 2026 08:09:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49919054</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49919054</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49919054</guid></item><item><title><![CDATA[New comment by croemer in "Backblaze drive stats for Q2 2026"]]></title><description><![CDATA[
<p>This seems like a poorly executed analysis of very interesting underlying data.<p>With this degree of average age variability they can't compare failure rates.<p>They should show us Kaplan-Meier (survival) curves not failure in this quarter.<p>Unless they set some datacenter on fire, not much in Q2 should be different vs Q1 for the same hard drive other than it being older.<p>Sure, internally, I'd do a better analysis controlled for age that shows if there are common failure rate increases but here they focus on model failure rates which is completely confounded by age differences.</p>
]]></description><pubDate>Wed, 30 Sep 2026 06:55:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49905348</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49905348</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49905348</guid></item><item><title><![CDATA[New comment by croemer in "America.gov"]]></title><description><![CDATA[
<p>For those outside US, in Europe I got "Service unavailable".<p>Switched on US VPN and it works.<p>Edit: Now it seems to work. Maybe it was a glitch. Or a US government agent is watching this page.</p>
]]></description><pubDate>Wed, 30 Sep 2026 06:11:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49905041</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49905041</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49905041</guid></item><item><title><![CDATA[New comment by croemer in "We found 24 Android vulnerabilities using our open source AI security agent"]]></title><description><![CDATA[
<p>The title is misleading, they found 24 vulnerabilities in Android _apps_ not in Android itself. That's a big difference.<p>In their own words from the body:<p>> I’ve reported more than 20 vulnerabilities in Android applications</p>
]]></description><pubDate>Tue, 29 Sep 2026 02:22:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887314</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49887314</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887314</guid></item><item><title><![CDATA[New comment by croemer in "Sonnet 5.5"]]></title><description><![CDATA[
<p>Pretty crazy that the model doesn't know that it needs to stop before it hits 128k output tokens. I guess it has no sense of how many tokens in it is? Wouldn't this be possible to work into the architecture?</p>
]]></description><pubDate>Mon, 28 Sep 2026 19:31:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49883198</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49883198</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49883198</guid></item><item><title><![CDATA[New comment by croemer in "Sonnet 5.5"]]></title><description><![CDATA[
<p>This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.</p>
]]></description><pubDate>Mon, 28 Sep 2026 18:52:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49882689</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49882689</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49882689</guid></item><item><title><![CDATA[New comment by croemer in "Sonnet 5.5"]]></title><description><![CDATA[
<p>Playing around with it for a few minutes, Sonnet 5.5 feels very fast, much quicker than Opus 5.5. Can't tell yet if it's a lot worse but the speed is definitely welcome.</p>
]]></description><pubDate>Mon, 28 Sep 2026 18:40:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49882502</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49882502</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49882502</guid></item><item><title><![CDATA[New comment by croemer in "Jev Can't Be Calibrated"]]></title><description><![CDATA[
<p>Ah, I see, you mean "Jev cannot possibly be calibrated for everyone as there is always missing context".<p>I understood it as "It is impossible to calibrate Jev".</p>
]]></description><pubDate>Sun, 27 Sep 2026 03:20:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49863056</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49863056</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49863056</guid></item><item><title><![CDATA[New comment by croemer in "Generate fonts where every LLM token is the same width"]]></title><description><![CDATA[
<p>I typed some German and it was breaking up words so much more than English. Not really surprising given tokenizers are optimized for most commonly used text.<p>Here's the token efficiency of a corpus translated into various languages and tokenized with the latest OpenAI one:<p><pre><code>  Language              Relative tokens
  --------------------------------------
  English                    1.00x
  Portuguese                 1.23x
  Chinese (Simplified)       1.25x
  German                     1.31x
  Spanish                    1.32x
  French                     1.37x
  Arabic                     1.38x
  Chinese (Traditional)      1.42x
  Korean                     1.47x
  Swahili                    1.49x
  Hindi                      1.57x
  Japanese                   1.66x
  Burmese                    3.16x
  Amharic                    5.78x
  Santali                   13.70x
</code></pre>
Source: "Tokenizer Fairness in 2026", a reproduction/extension of
Petrov, La Malfa, Torr & Bibi, "Language Model Tokenizers Introduce
Unfairness Between Languages" (NeurIPS 2023), using FLORES-200.<p><a href="https://github.com/partyfly/tokenizer-fairness-2026" rel="nofollow">https://github.com/partyfly/tokenizer-fairness-2026</a></p>
]]></description><pubDate>Sun, 27 Sep 2026 03:04:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49862972</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49862972</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49862972</guid></item><item><title><![CDATA[New comment by croemer in "Revealing the details of how OpenAI agents hacked Hugging Face"]]></title><description><![CDATA[
<p>TFA seems to be sloppy in writing, they should have kept the "meaning... [Some incorrect assumptions about GET]" out of the paragraph.</p>
]]></description><pubDate>Sat, 26 Sep 2026 10:02:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49854982</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49854982</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49854982</guid></item><item><title><![CDATA[New comment by croemer in "Revealing the details of how OpenAI agents hacked Hugging Face"]]></title><description><![CDATA[
<p>You're not missing anything. TFA really state this wrong assumption in their own voice.</p>
]]></description><pubDate>Sat, 26 Sep 2026 10:01:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49854967</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49854967</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49854967</guid></item><item><title><![CDATA[New comment by croemer in "Revealing the details of how OpenAI agents hacked Hugging Face"]]></title><description><![CDATA[
<p>> This access seems to have only allowed the agents to make ‘GET’ requests, meaning they could fetch and read websites, but not interact with them, submit forms, or send data to them<p>The authors of this (very interesting) analysis should really not state the sandbox's wrong assumptions in their own voice.<p>GET absolutely allows you to interact with sites. And of course GET can also send information. It's all up to the server that receives the GET to decide what it let's callers do with it.</p>
]]></description><pubDate>Sat, 26 Sep 2026 09:58:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49854953</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49854953</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49854953</guid></item><item><title><![CDATA[New comment by croemer in "Jev Can't Be Calibrated"]]></title><description><![CDATA[
<p>Article says that it _can_ be calibrated, it just isn't out of the box. So the title seems contradicted by the body.<p>> If you want calibrated probabilities you’ll still need to recalibrate Jev’s probabilities on your own data. The good news is that is cheap. A few hundred labeled examples from your actual data can be enough to fit a Platt scaling on top of Jev’s scores.</p>
]]></description><pubDate>Wed, 23 Sep 2026 18:22:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49820330</link><dc:creator>croemer</dc:creator><comments>https://news.ycombinator.com/item?id=49820330</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49820330</guid></item></channel></rss>