<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: sebnado</title><link>https://news.ycombinator.com/user?id=sebnado</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 22 Jul 2026 21:04:33 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=sebnado" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[Show HN: WorldBuild Bench repo: testing LLM world coherence with 3D games]]></title><description><![CDATA[
<p>I built WorldBuild Bench because, as we all know, llm bench scores often say something very different from what models actually feel like to use. It's really dependent on the type of tasks.<p>I personally want to test spatial, temporal, and causal coherence in an interactive 3D world. Does the model understand where things are, world stays consistent over time and do the consequences make sense?
There is a million people generating random games here and there on yt, but I want something that I can reproduce every time a model comes out and gets scored.<p>For this first run, 9 models received the same three roughly 30-line game prompt. I just added Kimi k3 to the results.<p>They all ran in high-thinking mode through the same open source harness, with the same sub-agent setup and access to Three.js, Rapier, and Playwright. There is currently one run per model per brief, producing 27 browser-playable games.<p>Because the qualities I’m interested in are difficult to score automatically, the main evaluation happens through blind pairwise comparisons. You play two games built from the same brief without seeing the model names, then compare their game feel, world design, presentation, completeness, and overall quality.<p>I’m also publishing the prompts, generated artifacts, generation time, estimated cost, and code size.<p>Fable produced some of the strongest games from what I could see, but its three runs cost about $756. GPT-5.6 Sol cost about $108, while GLM-5.2 and Grok 4.5 each cost around $19. Opus also felt closer to Fable than I expected, considering the large cost difference.
Kimi somehow ended up roughly at the cost of GPT, but performed somewhat similar to Opus (thats just my subjective opinion there)<p>This first run is small, and as stated above, human preference is subjective. But I plan on running more and hope to evolve the methodology. As long as I can afford all these tokens. Fable is ridiculously expensive.<p>If you look at it, I'd really appreciate criticism of the task methodology, design, blind evaluation, etc. What would make this rigorous enough to be truly useful.</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48994848">https://news.ycombinator.com/item?id=48994848</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 21 Jul 2026 16:48:39 +0000</pubDate><link>https://github.com/sebnado/worldbuild-bench</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=48994848</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48994848</guid></item><item><title><![CDATA[New comment by sebnado in "Show HN: Wheesper – Start an anonymous discussion with a link"]]></title><description><![CDATA[
<p>I think it would benefit from having a upvote type of feature to surface opinions/comments that people agree with. if you have many people with the same insight in your thread, there is no point in restating the same point over and over to show you agree</p>
]]></description><pubDate>Mon, 20 Jul 2026 18:01:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48982451</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=48982451</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48982451</guid></item><item><title><![CDATA[New comment by sebnado in "Jelly UI: Soft-body physics for native HTML form controls"]]></title><description><![CDATA[
<p>love it I can think of many projects that could use that</p>
]]></description><pubDate>Mon, 20 Jul 2026 17:54:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48982307</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=48982307</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48982307</guid></item><item><title><![CDATA[New comment by sebnado in "Ask HN: What will humans do when AI writes and reviews code?"]]></title><description><![CDATA[
<p>We’ve been running an AI-first dev loop in production for ~2 years (disclaimer: I help build Ze1 and Sandscape, they are both Ai driven products). A few things we’ve learned:<p>Instead of cranking boilerplate, they spend Day 1 reviewing AI diffs. We pair them with a senior for a “why did the agent choose this?” teardown.<p>What matters is mean-time-to-rollback. If the agent + test harnesses catch breakage faster than a human pair can, 2 × better is already good economics. Reliability engineering beats perfection thresholds.<p>Syntax, style and most unit-level bugs are now linted or auto-fixed. Humans zoom out to architecture, data contracts, threat models, perf budgets, and “does this change make sense for the product?”. So even juniors are now a lot more involved on the sujective elements of development<p>So I think that as the easy bugs vanish, new failure modes show up: latency cliffs, subtle privacy leaks, energy use, fairness. The goal posts moves</p>
]]></description><pubDate>Mon, 02 Jun 2025 18:14:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=44161523</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=44161523</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44161523</guid></item><item><title><![CDATA[New comment by sebnado in "Is AI sparking a cognitive revolution leading to mediocrity and conformity?"]]></title><description><![CDATA[
<p>Fascinating read. I see the same tension: today’s models are masters of interpolation, but they still lack intentionality. If we push them toward “cognitive exoskeletons” (handling the combinatorial grunt work while humans dictate the why) we get amplification instead of flattening.</p>
]]></description><pubDate>Mon, 02 Jun 2025 18:03:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=44161438</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=44161438</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44161438</guid></item><item><title><![CDATA[New comment by sebnado in "AI Is Learning to Escape Human Control"]]></title><description><![CDATA[
<p>Thing whole thing is a bit dumb tbh. You send conflicting requests to the LLM and it fails at doing both. It's nothing new, we all know it. Every article's headline make it sound like the AI is somehow consciously refusing the request to shutdown, while the only thing it demonstrate is that we are over fitting for specific outcomes. We are still a long way from human level general intelligence.</p>
]]></description><pubDate>Mon, 02 Jun 2025 17:57:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=44161385</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=44161385</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44161385</guid></item><item><title><![CDATA[New comment by sebnado in "Claude Code: An Agentic cleanroom analysis"]]></title><description><![CDATA[
<p>By saying that's its gold mine, I think OP meant that's it's funny, not that it brings valuable insight. 
ie: THEY KNOW -> that made me laugh<p>and as the article said 
"an LLM who just spent thousands of words explaining why they're not allowed to use thousands of words", its just funny to read.</p>
]]></description><pubDate>Sun, 01 Jun 2025 23:40:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=44154608</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=44154608</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44154608</guid></item><item><title><![CDATA[New comment by sebnado in "Making maps with noise functions (2022)"]]></title><description><![CDATA[
<p>It's a really good approach for procedural maps, but if you want to achieve realism, you'll likely need something like Gaea</p>
]]></description><pubDate>Sun, 01 Jun 2025 23:35:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=44154581</link><dc:creator>sebnado</dc:creator><comments>https://news.ycombinator.com/item?id=44154581</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44154581</guid></item></channel></rss>