<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Davisb135</title><link>https://news.ycombinator.com/user?id=Davisb135</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 28 Jul 2026 12:36:17 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Davisb135" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by Davisb135 in "The case for MUDs in modern times (2018)"]]></title><description><![CDATA[
<p>I'm glad it sparked something. Really cool to see not just the nostalgia, but the active projects/games, and the feature requests. Original Post: <a href="https://news.ycombinator.com/item?id=49008538">https://news.ycombinator.com/item?id=49008538</a></p>
]]></description><pubDate>Sat, 25 Jul 2026 03:20:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49044197</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49044197</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49044197</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>Just starting small. Those are all ideas for Phases 2 & 3.</p>
]]></description><pubDate>Thu, 23 Jul 2026 01:35:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49015751</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49015751</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49015751</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>Haha, thanks. I wish I'd come up with that. It's from Gunpei Yokoi of Nintendo fame.</p>
]]></description><pubDate>Wed, 22 Jul 2026 20:49:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49013224</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49013224</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49013224</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>Having stepped away from MUDs ourselves for years, it's been interesting to see what the community has been cooking up in terms of UI/UX, Evennia, MUDlet, TUIs, etc.</p>
]]></description><pubDate>Wed, 22 Jul 2026 18:59:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49011738</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49011738</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49011738</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>That's fair. We leaned on LLMs to help draft parts of the paper but disclosed it in the acknowledgements section. We didn't appreciate the amount of work this "hobby" experiment would require, and so looked for ways to help get this out the door.</p>
]]></description><pubDate>Wed, 22 Jul 2026 18:57:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=49011716</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49011716</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49011716</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>That's a cool project, thanks for sharing!</p>
]]></description><pubDate>Wed, 22 Jul 2026 18:53:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49011653</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49011653</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49011653</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>Not Hitchhiker's specifically, we'd considered another existing one but as soon as we started digging into IP law related to MUDs, we decided if we're going to do it, we'll just start our own from scratch.</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:53:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010720</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49010720</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010720</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>That's a good idea, thanks! We'd only considered the web app + using a client.</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:45:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010593</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49010593</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010593</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>Thanks! Yeah, it started with a friend's idea to take a MUD we played in the 90s / early 2000s that had been open sourced and adapt it to mobile games. He's a middle school teacher and wanted to introduce the next generations to MUDs.<p>We then decided to pivot and thought, what if this would be a good eval environment for LLMs from the simplistic standpoint of, text is their native environment.</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:29:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010320</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49010320</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010320</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>Sadly, you're not wrong... It might not get off the ground, but if we're able to, we'll give it a shot.</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:25:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010263</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49010263</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010263</guid></item><item><title><![CDATA[New comment by Davisb135 in "Can a MUD evaluate LLMs? A $99 proof of concept"]]></title><description><![CDATA[
<p>I agree with you. You've also got more / easier competition. IIRC, MUDs peaked in the mid-late 90s. MMOs were really the death knell, then the explosion of other types of games, mobile games, then to your point things like Discord.<p>One adjacent project we're exploring for the future is if you took all of the current understanding of game design and mechanics with modern AI functionality, knowledge bases, etc, and built a MUD, what would it look like and would it be enough to bring back players from other games?</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:18:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010133</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49010133</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010133</guid></item><item><title><![CDATA[Can a MUD evaluate LLMs? A $99 proof of concept]]></title><description><![CDATA[
<p>I'm the author of a paper my friends and I wrote after we were curious if a MUD, text games originating in the 1970s, could be used to evaluate LLMs. We've spent the last several months on nights and weekends running this experiment and writing the paper on just our personal computers with about $99 in API credits.<p>Our experiment did have an interesting leaderboard but even more surprising was the measurements of each LLM. We scored each on four behavioral dimensions, two of which lean heavily on an LLM classifier. When we removed those two, one of the frontier models fell six positions. When we then checked the classifier against a second judge, the per-model agreement between them ranged from 85% to 22%. The aggregate kappa (0.04 on probe detection) indicated the instrument was noisy without saying which models the noise was hitting. The most affected model shared a model family with the classifier. This isn't proof of bias, just one observation we recorded.<p>We realize LLM judges can be unreliable, and while it wasn't our original intent to test this, it ended up being the most interesting finding. The divergence between the two judges is the finding we think generalizes to other judge-based benchmarks.<p>We emphasize this is just a proof of concept and not a validated benchmark. We prepared a thorough limitations section in the paper, including just 50 runs per model, overlapping CIs among the top models, no human raters, a tiny environment, etc.<p>Everything we did is publicly available, the paper and data are CC BY 4.0, while the code is MIT. The paper, transcripts, code, and complete API billing export  can be found at <a href="https://doi.org/10.5281/zenodo.21386663" rel="nofollow">https://doi.org/10.5281/zenodo.21386663</a><p>If you find issues, please let us know, that's why we're sharing this. We're currently designing Phase 2 and want it to be as robust as possible. We're looking at human baselines, multiple judges, more objectives, a larger environment, etc.</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49008538">https://news.ycombinator.com/item?id=49008538</a></p>
<p>Points: 109</p>
<p># Comments: 78</p>
]]></description><pubDate>Wed, 22 Jul 2026 15:39:01 +0000</pubDate><link>https://cruciblebench.ai/</link><dc:creator>Davisb135</dc:creator><comments>https://news.ycombinator.com/item?id=49008538</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49008538</guid></item></channel></rss>