<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: lonlundgren</title><link>https://news.ycombinator.com/user?id=lonlundgren</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 22 Sep 2026 00:52:09 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=lonlundgren" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by lonlundgren in "Fable 5 – Median thinking declined in August"]]></title><description><![CDATA[
<p>After reading what you had to add to this discussion, I have only two things to leave you with:<p>1/ if the point of setting xhigh or max effort on a "reasoning model" is -not- to have additional reasoning tokens applied to the problem, then I fear we are all using AI wrong
2/ the decorum with which you "review" someone's work in public is entirely your choice. the only "cheap trick" on the internet is being a pseudonymous ass.</p>
]]></description><pubDate>Mon, 21 Sep 2026 19:24:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49792084</link><dc:creator>lonlundgren</dc:creator><comments>https://news.ycombinator.com/item?id=49792084</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49792084</guid></item><item><title><![CDATA[New comment by lonlundgren in "Fable 5 – Median thinking declined in August"]]></title><description><![CDATA[
<p>Thank you for your kind words, Aurornis.<p>Yes, this is my production corpus, across 65 usage days, two subscription accounts, three machines, 25 project groups, and 213 sessions, across 43,261 invocations and 7,583 turns. Use your own data if you want to prove or refute what was seen in my corpus.<p>The "ups and downs in the chart" were not plotted with sub-daily resolution. Specifically, the two-month temporal chart uses a 3.5-day Gaussian bandwidth which is meant to reduce short-term noise while retaining broader changes.<p>Additionally, a separate episodic analysis identified multi-day changes in delivered thinking. And those episodes were predictive of held-out work. The more projects pulled into an ensemble, the more predictive they were of delivered thinking tokens for held-out projects during the episode.<p>The point of the benchmarks is to establish a baseline for what thinking-token counts one should expect from specific effort levels using published numbers, not what every response should be delivered. If you looked closer, you would see that P90 invocations were still delivered 13x thinking tokens below that level.<p>So you don't have to provide a generous interpretation of my workload if you don't want to. Remove all of the zero-thinking token responses, redistribute those samples across the distribution, and then tell me if it magically shifts right and starts delivering anything close to published numbers. Only -46- out of 36,374 July and August invocations even broke 16k thinking tokens - only 0.13% of the total. You are welcome to present what percentile you think is a fair comparison to make here.<p>If you would like to denigrate a month of my time as vibe-slop, that is your prerogative. You can even be dismissive of my workload, if you want, even if my background should tell you otherwise. But if you want to knock a month of someone's time, do it with your own data to at least help move the conversation forward.</p>
]]></description><pubDate>Mon, 21 Sep 2026 19:08:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49791866</link><dc:creator>lonlundgren</dc:creator><comments>https://news.ycombinator.com/item?id=49791866</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49791866</guid></item></channel></rss>