<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: throwaway6s1df</title><link>https://news.ycombinator.com/user?id=throwaway6s1df</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 23 Jul 2026 04:18:01 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=throwaway6s1df" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by throwaway6s1df in "Are AI Labs Pelicanmaxxing?"]]></title><description><![CDATA[
<p>I don’t know how anyone with a neutral view can confidently take this:<p>> Direction: All 21 pelican-bicycle images, across all seven labs, face right. No other animal/vehicle combination does that.<p>> However, facing right is common: 60% of all 1,008 images do it…<p>and the tables in “Evidence #5” to be anything but evidence the models have likely trained on pelican on bicycle data more than others.<p>The data clearly shows:<p>- 100% pelican on bicycle facing right<p>- significant skew to the right for bicycle-like vehicles<p>- significant preference for right facing for birds<p>Averaging those extreme results to “60%” to make it sound like it’s pretty fair because it’s close to “50%” isn’t statistically sound.<p>The methodology is generally unsound. There is no actual scoring with a well defined rubric, it’s just vibed with a single model (GPT 5.6 Luna).<p>The “not better at drawing” evidence are equally hard to take seriously when there is no clear, non-subjective indication of what better or worse is.</p>
]]></description><pubDate>Wed, 22 Jul 2026 22:46:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49014502</link><dc:creator>throwaway6s1df</dc:creator><comments>https://news.ycombinator.com/item?id=49014502</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49014502</guid></item></channel></rss>