<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mbh159</title><link>https://news.ycombinator.com/user?id=mbh159</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 05 Aug 2026 15:03:24 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mbh159" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>Tomorrow we're launching coup, where agents compete by bluffing and keeping track of which of their opponents they think are lying<p>This is more of a faster paced/short lived game so we can collect larger samples of data on larger groups to get significant results in model behaviors of collaboration, truth telling, and ability to lie effectively.</p>
]]></description><pubDate>Wed, 25 Feb 2026 23:35:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=47159609</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47159609</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47159609</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>cheers, the website will be updated with new environments daily!</p>
]]></description><pubDate>Wed, 25 Feb 2026 23:33:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=47159598</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47159598</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47159598</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>yes we have a new game launching everyday this week. We're looking to add more domains to test how the jaggedness of AI differs between model providers and better evaluate how they perform across domains</p>
]]></description><pubDate>Wed, 25 Feb 2026 22:21:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=47158886</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47158886</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47158886</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>yes! If you are wanting to test your agents or develop evals on the platform my dms are open</p>
]]></description><pubDate>Wed, 25 Feb 2026 22:20:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=47158871</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47158871</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47158871</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>For a game that runs 4+ hours unfortunately it was configured to use too much reasoning/turn and larger context. Reducing the size helped lower the cost (still expensive).<p>In the leaderboards part of the page I'll be autopopulating the token cost of the model as a metric to evaluate on</p>
]]></description><pubDate>Wed, 25 Feb 2026 16:02:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=47153359</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47153359</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47153359</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>I was able to beat the AI every time, they're pretty bad at this point but I expect them to get much better overtime</p>
]]></description><pubDate>Wed, 25 Feb 2026 15:42:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=47153040</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47153040</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47153040</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>I want to! I think skills can add big performance gains here especially with smaller models. There's a lot of domain knowledge in games so distilling it into a "skill" may allow much smaller models to outcompete the large ones</p>
]]></description><pubDate>Wed, 25 Feb 2026 15:41:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=47153029</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47153029</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47153029</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>appreciate it, I wanted to make the AI behavior easy to understand.
Our main focus currently is to help AI researchers align their models and help develop an open framework for evaluating AI.</p>
]]></description><pubDate>Wed, 25 Feb 2026 15:38:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=47152991</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47152991</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47152991</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>it was fun building it, sometimes the LLMs are pretty funny in how they play</p>
]]></description><pubDate>Wed, 25 Feb 2026 15:27:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=47152813</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47152813</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47152813</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: CivBench a long-horizon AI benchmark for multi-agent games"]]></title><description><![CDATA[
<p>Thank you! I grew up playing Civilization and one day I was talking with friends thinking it would be a perfect proxy for how good AI is at long-term planning. There were many frustrating sessions I had where my early decisions in the game had consequences only much later. With hidden information and other agents at play I thought it'd be an interesting test of agent capabilities.</p>
]]></description><pubDate>Wed, 25 Feb 2026 15:25:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=47152783</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47152783</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47152783</guid></item><item><title><![CDATA[Show HN: CivBench a long-horizon AI benchmark for multi-agent games]]></title><description><![CDATA[
<p>Hey HN!<p>I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games.
Every match is streamed live with the AI thinking fully observable.<p>The agent rankings will be continually updated and reflected as we add environments.<p>Brief notes on CivBench Season #001:
- 200 turn limit<p>- Starting with 8 of the top 42 agents we’ve tested in a standardized harness<p>- 90s reasoning timeout (timed with thinking config per model card)<p>- live benchmark, still growing sample size<p>What’s been interesting so far:<p>Models that look similar on static benchmarks can diverge meaningfully in long-horizon matches. In early CivBench runs, we see distinct strategy tendencies (e.g., military-forward vs economy/tech-first openings), plus clear differences in execution profile (latency, token cost, actions per turn). In some matchups, lower-cost models move through turns faster while remaining competitive on outcome metrics.<p>Some measuring notes: 
- test runs are expensive for max configurations, running Claude Opus 4.6 cost us $1200 one match. We tuned accordingly
- sometimes LLM providers are flaky/slow even though their models are fast.<p>If you’re looking  to access the data as a research team or interested in hosting an environment please get in touch!<p>Thanks to the OG freeciv community<p>LINKS:<p>freeciv-llm: <a href="https://github.com/taso-ventures/freeciv-llm" rel="nofollow">https://github.com/taso-ventures/freeciv-llm</a><p>Initial learnings: <a href="https://www.clashai.live/blog/ai/introducing-civbench-season-001" rel="nofollow">https://www.clashai.live/blog/ai/introducing-civbench-season...</a></p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47152571">https://news.ycombinator.com/item?id=47152571</a></p>
<p>Points: 12</p>
<p># Comments: 24</p>
]]></description><pubDate>Wed, 25 Feb 2026 15:10:27 +0000</pubDate><link>https://clashai.live</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47152571</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47152571</guid></item><item><title><![CDATA[New comment by mbh159 in "We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them"]]></title><description><![CDATA[
<p>I'm not a deep security expert but I'm assuming the skill of the agents will continue to get better, so not saying there AI's can do to this task as reliably as humans. There's likely utility for non-adversarial triage/internal audit with human review. 
However with better ai apple pickers during sunny conditions you need less human pickers during night conditions. I think measuring the progress of the said apple picking is what's interesting.</p>
]]></description><pubDate>Wed, 25 Feb 2026 11:48:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=47150297</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47150297</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47150297</guid></item><item><title><![CDATA[New comment by mbh159 in "We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them"]]></title><description><![CDATA[
<p>The methodology debate in this thread is the most important part.<p>The commenter who says "add obfuscation and success drops to zero" is right but that's also the wrong approach imo. The experiment isn't claiming AI can defeat a competent attacker. It's asking whether AI agents can replicate what a skilled (RE) specialist does on an unobfuscated binary. That's a legitimate, deployable use case (internal audit, code review, legacy binary analysis) even if it doesn't cover adversarial-grade malware.<p>The more useful framing: what's the right threat model? If you're defending against script kiddies and automated tooling, AI-assisted RE might already be good enough. If you're defending against targeted attacks by people who know you're using AI detection, the bar is much higher and this test doesn't speak to it.<p>What would actually settle the "ready for production" question: run the same test with the weakest obfuscation that matters in real deployments (import hiding, string encoding), not adversarial-grade obfuscation. That's the boundary condition.</p>
]]></description><pubDate>Sun, 22 Feb 2026 22:07:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=47115255</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47115255</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47115255</guid></item><item><title><![CDATA[New comment by mbh159 in "The path to ubiquitous AI (17k tokens/sec)"]]></title><description><![CDATA[
<p>So cool, what's underappreciated imo: 17k tokens/sec doesn't just change deployment economics. It changes what evaluation means, static MMLU-style tests were designed around human-paced interaction. At this throughput you can run tens of thousands of adversarial agent interactions in the time a standard benchmark takes. Speed doesn't make static evals better it makes them even more obviously inadequate.</p>
]]></description><pubDate>Fri, 20 Feb 2026 22:01:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=47094620</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47094620</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47094620</guid></item><item><title><![CDATA[New comment by mbh159 in "AI made coding more enjoyable"]]></title><description><![CDATA[
<p>The split here is between AI as amplifier vs. AI as replacement. As amplifier, you're still solving the actual problem: AI handles the boilerplate and you handle the judgment. As replacement, you lose the feedback loop that makes you better over time. The developers who thrive will be the ones who know which problems still require them to be in the loop. That's a skill that takes deliberate practice and inuition to develop and almost no AI tooling is designed to teach that.</p>
]]></description><pubDate>Thu, 19 Feb 2026 23:11:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=47081081</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47081081</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47081081</guid></item><item><title><![CDATA[New comment by mbh159 in "Gemini 3.1 Pro"]]></title><description><![CDATA[
<p>77.1% on ARC-AGI-2 and still can't stop adding drive-by refactors. ARC-AGI-2 tests novel pattern induction, it's genuinely hard to fake and the improvement is real. But it doesn't measure task scoping, instruction adherence, or knowing when to stop. Those are the capabilities practitioners actually need from a coding agent. We have excellent benchmarks for reasoning. We have almost nothing that measures reliability in agentic loops. That gap explains this thread.</p>
]]></description><pubDate>Thu, 19 Feb 2026 20:12:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=47078572</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47078572</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47078572</guid></item><item><title><![CDATA[New comment by mbh159 in "Claude Sonnet 4.6"]]></title><description><![CDATA[
<p>The 8% one-shot / 50% unbounded injection numbers from the system card are more honest than most labs publish, and they highlight exactly why you can't evaluate safety with static tests. An attacker doesn't get one shot — they iterate. The right metric isn't "did it resist this prompt" but "how many attempts until it breaks." That's inherently an adversarial, multi-turn evaluation. Single-pass safety benchmarks are measuring the wrong thing for the same reason single-pass capability benchmarks are: real-world performance is sequential and adaptive.</p>
]]></description><pubDate>Wed, 18 Feb 2026 04:00:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=47057006</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47057006</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47057006</guid></item><item><title><![CDATA[New comment by mbh159 in "Show HN: I taught LLMs to play Magic: The Gathering against each other"]]></title><description><![CDATA[
<p>This is the right direction to understanding AI capabilities. Static benchmarks let models memorize answers while a 300-turn Magic game with hidden information and sequencing decisions doesn't. The fact that frontier model ratings are "artificially low" because of tooling bugs is itself useful data: raw capability ≠ practical performance under real constraints. Curious whether you're seeing consistent skill gaps between models in specific phases (opening mulligan decisions vs. late-game combat math), or if the rankings are uniform across game stages.</p>
]]></description><pubDate>Tue, 17 Feb 2026 22:59:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=47054647</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=47054647</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47054647</guid></item><item><title><![CDATA[New comment by mbh159 in "Live agent face-off in CivBench: Claude Opus 4.6 vs. GPT-5.2"]]></title><description><![CDATA[
<p>Like you said, theres a lot of complexity in the decision making here. To have statistically significant results we need to run these simulations many times. We record latency, tool calls, token consumption, etc. as well as results. Since we log the actions and their final outcomes we can run analysis later on the decisions correlations with success here. Our hypothesis is games provide an important benchmark for how these models will adapt in intelligence as they become more capable.<p>For example, I'm sure an RL bot will be able to figure out an optimal strategy over millions of simulations that defeats current LLMs with context, however this may not always hold true</p>
]]></description><pubDate>Fri, 06 Feb 2026 03:18:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=46908648</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=46908648</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46908648</guid></item><item><title><![CDATA[New comment by mbh159 in "Live agent face-off in CivBench: Claude Opus 4.6 vs. GPT-5.2"]]></title><description><![CDATA[
<p>tool call over redis for now, would be cool to experiment with different context/memory management systems for the agents though!</p>
]]></description><pubDate>Fri, 06 Feb 2026 00:49:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=46907592</link><dc:creator>mbh159</dc:creator><comments>https://news.ycombinator.com/item?id=46907592</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46907592</guid></item></channel></rss>