<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: meander_water</title><link>https://news.ycombinator.com/user?id=meander_water</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 28 Jul 2026 12:20:05 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=meander_water" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by meander_water in "Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard"]]></title><description><![CDATA[
<p>Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate.<p>> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.<p>This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.</p>
]]></description><pubDate>Sat, 25 Jul 2026 14:09:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49047697</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=49047697</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49047697</guid></item><item><title><![CDATA[New comment by meander_water in "Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard"]]></title><description><![CDATA[
<p>The funny thing is that these leaderboards have become completely meaningless for end-users to make decisions on when to use what model.<p>A single metric ranking is useless because each model has strengths and weaknesses for specific domains and tasks. There is no "one best model" anymore, and you might not even need the best model for the level of complexity for your task.<p>For e.g. you might use Fable for UI design, Sol for systems design backend work and Kimi K3 for exploit development.<p>The only purpose these metrics serve is bragging rights for the model companies.</p>
]]></description><pubDate>Sat, 25 Jul 2026 10:19:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49046304</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=49046304</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49046304</guid></item><item><title><![CDATA[Substack adds AI text detection to all notes and posts]]></title><description><![CDATA[
<p>Article URL: <a href="https://post.substack.com/p/against-claudefishing">https://post.substack.com/p/against-claudefishing</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49045198">https://news.ycombinator.com/item?id=49045198</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 25 Jul 2026 07:04:43 +0000</pubDate><link>https://post.substack.com/p/against-claudefishing</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=49045198</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49045198</guid></item><item><title><![CDATA[New comment by meander_water in "OpenAI’s accidental attack against Hugging Face is science fiction that happened"]]></title><description><![CDATA[
<p>1. The benchmark is run with a python script - <a href="https://github.com/sunblaze-ucb/exploitgym" rel="nofollow">https://github.com/sunblaze-ucb/exploitgym</a> using an agent harness. I suspect they used codex. So the model has access to the environment and could trivially inspect its own source code, which has lots of references to exploit gym and docs relating to it.<p>2. Hanlons Razor</p>
]]></description><pubDate>Thu, 23 Jul 2026 22:36:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49028976</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=49028976</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49028976</guid></item><item><title><![CDATA[New comment by meander_water in "Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models"]]></title><description><![CDATA[
<p>Seems similar to Openrouter Fusion - <a href="https://openrouter.ai/docs/guides/routing/routers/fusion-router" rel="nofollow">https://openrouter.ai/docs/guides/routing/routers/fusion-rou...</a></p>
]]></description><pubDate>Thu, 23 Jul 2026 21:25:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49028256</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=49028256</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49028256</guid></item><item><title><![CDATA[AI Compass: which archetype are you?]]></title><description><![CDATA[
<p>Article URL: <a href="https://bambamramfan.github.io/ai-compass/">https://bambamramfan.github.io/ai-compass/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48732217">https://news.ycombinator.com/item?id=48732217</a></p>
<p>Points: 9</p>
<p># Comments: 1</p>
]]></description><pubDate>Tue, 30 Jun 2026 13:07:01 +0000</pubDate><link>https://bambamramfan.github.io/ai-compass/</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48732217</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48732217</guid></item><item><title><![CDATA[Meta Posed as Teens to Prompt Rival Chatbots About Suicide, Sex, and Drugs]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.wired.com/story/meta-contractors-pretending-to-be-teens-chatbot-testing/">https://www.wired.com/story/meta-contractors-pretending-to-be-teens-chatbot-testing/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48726377">https://news.ycombinator.com/item?id=48726377</a></p>
<p>Points: 28</p>
<p># Comments: 8</p>
]]></description><pubDate>Mon, 29 Jun 2026 22:51:48 +0000</pubDate><link>https://www.wired.com/story/meta-contractors-pretending-to-be-teens-chatbot-testing/</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48726377</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48726377</guid></item><item><title><![CDATA[New comment by meander_water in "Micro-Agent: Beat Frontier Models with Collaboration Inside Model API"]]></title><description><![CDATA[
<p>I thought all model providers are doing this under the hood anyway in their UI?<p>They certainly seem to when A/B testing different models, and Fable routes to Opus 4.8 when guardrails fail.<p>Also, openrouter recently released a fusion router - <a href="https://openrouter.ai/blog/announcements/fusion-beats-frontier/" rel="nofollow">https://openrouter.ai/blog/announcements/fusion-beats-fronti...</a></p>
]]></description><pubDate>Mon, 29 Jun 2026 21:50:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48725766</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48725766</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48725766</guid></item><item><title><![CDATA[New comment by meander_water in "Exploring the internal representations of Pangram 3.3.2"]]></title><description><![CDATA[
<p>GPTZero is much better at handling humanized outputs. Also has a similar false positive rate to Pangram.</p>
]]></description><pubDate>Thu, 25 Jun 2026 04:03:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=48668729</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48668729</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48668729</guid></item><item><title><![CDATA[New comment by meander_water in "Stealing Is a Skill"]]></title><description><![CDATA[
<p>> However, it’s your job to go down the rabbit hole, learn the 100%, and sprinkle in your 3%.<p>I would say that there is a big difference between stealing without acknowledgement, and stealing with acknowledgement and actively learning through reverse engineering.</p>
]]></description><pubDate>Wed, 24 Jun 2026 14:16:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48660300</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48660300</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48660300</guid></item><item><title><![CDATA[New comment by meander_water in "Use AI for reviewing code especially when the diff is huge"]]></title><description><![CDATA[
<p>> I don't think you should waste time reviewing every single line of code in here and just use AI to review it!<p>> What you bring is the knowledge that the author nor the LLM doesn't know.<p>How can you possibly know what relevant context to provide the LLM unless you read the 10k loc? Now you've wasted double the time.</p>
]]></description><pubDate>Mon, 22 Jun 2026 10:40:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48628364</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48628364</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48628364</guid></item><item><title><![CDATA[New comment by meander_water in "GLM 5.2 vs. Opus"]]></title><description><![CDATA[
<p>Thanks, I didn't mean to be brusque, but I have seen a lot of these vibe tests lately that come to grand conclusions like "X model is better than Y" from the result of a single prompt.<p>Appreciate you sharing the results of your tests though!</p>
]]></description><pubDate>Mon, 22 Jun 2026 08:01:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48627163</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48627163</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48627163</guid></item><item><title><![CDATA[New comment by meander_water in "GLM 5.2 vs. Opus"]]></title><description><![CDATA[
<p>> So we ran it head-to-head against Claude Opus 4.8: same one-shot prompt, build a 3D platformer in raw WebGL from scratch<p>Running a single one-shot prompt is not a benchmark, not is it representative of any sort of real-world usage.<p>Most agent usage is collaborative so you need to test things like reliability (when I delegate a task, does it complete it without making up test results for e.g.) and steerability (does it obey my instructions or does it just do what it thinks is best).</p>
]]></description><pubDate>Mon, 22 Jun 2026 07:42:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48627015</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48627015</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48627015</guid></item><item><title><![CDATA[Machine Studying]]></title><description><![CDATA[
<p>Article URL: <a href="https://jacobxli.com/blog/2026/machine-studying/">https://jacobxli.com/blog/2026/machine-studying/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48623669">https://news.ycombinator.com/item?id=48623669</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 21 Jun 2026 23:26:12 +0000</pubDate><link>https://jacobxli.com/blog/2026/machine-studying/</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48623669</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48623669</guid></item><item><title><![CDATA[New comment by meander_water in "Building an HTML-first site doubled our users overnight"]]></title><description><![CDATA[
<p>Really curious to understand why I'm being downvoted. I don't think it's a particularly spicy take - Just choose the right tool for the job.</p>
]]></description><pubDate>Thu, 11 Jun 2026 01:44:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48485287</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48485287</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48485287</guid></item><item><title><![CDATA[New comment by meander_water in "Building an HTML-first site doubled our users overnight"]]></title><description><![CDATA[
<p>As someone who has built both react based frontends and html based ones (with htmx), there is a law of diminishing returns at play.<p>To start off, writing a basic crud website with forms is much easier with htmx.<p>But when you start building more complex components, and integrate with other systems (OAuth for e.g.)  there are tons of libraries and SDKs for the react ecosystem, but not many for pure html components.<p>At this point, it's much easier to use off the shelf components than it is to manually write html to handle all the bizarre UI edge cases.</p>
]]></description><pubDate>Wed, 10 Jun 2026 22:10:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48483411</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48483411</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48483411</guid></item><item><title><![CDATA[New comment by meander_water in "Claude Fable 5"]]></title><description><![CDATA[
<p>All the model releases we've seen this year have only made incremental improvements in benchmarks.<p>This feels like the first release that feels like a significant step up in terms of benchmark results.<p>Can anyone make an educated guess what the secret sauce in the model architecture is between 4.8 and Fable?</p>
]]></description><pubDate>Tue, 09 Jun 2026 23:08:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48469073</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48469073</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48469073</guid></item><item><title><![CDATA[What a Joke: GitHub Copilots Token Based Billing Spurs Consternation]]></title><description><![CDATA[
<p>Article URL: <a href="https://techcrunch.com/2026/05/30/what-a-joke-github-copilots-new-token-based-billing-spurs-consternation-among-devs/">https://techcrunch.com/2026/05/30/what-a-joke-github-copilots-new-token-based-billing-spurs-consternation-among-devs/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48355911">https://news.ycombinator.com/item?id=48355911</a></p>
<p>Points: 3</p>
<p># Comments: 1</p>
]]></description><pubDate>Mon, 01 Jun 2026 12:25:41 +0000</pubDate><link>https://techcrunch.com/2026/05/30/what-a-joke-github-copilots-new-token-based-billing-spurs-consternation-among-devs/</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48355911</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48355911</guid></item><item><title><![CDATA[New comment by meander_water in "The four-day workweek in Australia: insights from early adopters of 100:80:100"]]></title><description><![CDATA[
<p>Not the first study, and they all largely report the same results:<p><a href="https://www.nature.com/articles/s41562-025-02259-6" rel="nofollow">https://www.nature.com/articles/s41562-025-02259-6</a><p><a href="https://www.theguardian.com/money/2019/feb/19/four-day-week-trial-study-finds-lower-stress-but-no-cut-in-output" rel="nofollow">https://www.theguardian.com/money/2019/feb/19/four-day-week-...</a><p><a href="https://www.4dayweek.com/research" rel="nofollow">https://www.4dayweek.com/research</a></p>
]]></description><pubDate>Sun, 24 May 2026 23:10:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=48261932</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48261932</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48261932</guid></item><item><title><![CDATA[New comment by meander_water in "I Miss Terry Pratchett"]]></title><description><![CDATA[
<p>Lovely sentiment in the article, which was unfortunately AI generated.<p>Can we start tagging titles in HN with [AI-generated] or something?<p>I know some people have no problem with it, but it might help others (like me) to steer clear</p>
]]></description><pubDate>Sat, 23 May 2026 13:07:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48247346</link><dc:creator>meander_water</dc:creator><comments>https://news.ycombinator.com/item?id=48247346</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48247346</guid></item></channel></rss>