<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: eis</title><link>https://news.ycombinator.com/user?id=eis</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 06 Sep 2026 18:19:26 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=eis" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by eis in "GPT-6 Astra makes major gains in the Artificial Analysis Coding Agent Index"]]></title><description><![CDATA[
<p>In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58).
In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.<p>Am I missing something or is this not looking too... stellar?</p>
]]></description><pubDate>Thu, 03 Sep 2026 21:09:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49557052</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49557052</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49557052</guid></item><item><title><![CDATA[New comment by eis in "Elevated Errors for Multiple Models"]]></title><description><![CDATA[
<p>That's a good point. Seems like Bedrock offers the same pricing while also providing an uptime SLA.</p>
]]></description><pubDate>Thu, 03 Sep 2026 15:56:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49552139</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49552139</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49552139</guid></item><item><title><![CDATA[New comment by eis in "Elevated Errors for Multiple Models"]]></title><description><![CDATA[
<p>Cost and reliability are the two reasons why we don't use Claude in our product. Getting close to one nine, that's not something one can build a reliable product upon. We now use OpenAI with Gemini fallback (or vice versa depending on use case).
Personally I like Claude and have the 20x Max plan but even there I burned through the whole weekly quota with 3 prompts in less than a day using the new Fable 5.1 which is crazy. Now Opus 5 is down. These two issues are really testing my patience.</p>
]]></description><pubDate>Thu, 03 Sep 2026 13:58:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49550007</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49550007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49550007</guid></item><item><title><![CDATA[New comment by eis in "Gemini 3.8 Flash and 3.8 Flash Cyber"]]></title><description><![CDATA[
<p>3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just increased the thinking budgets...<p>3.7 used 64M on high: <a href="https://artificialanalysis.ai/models/gemini-3-7-flash" rel="nofollow">https://artificialanalysis.ai/models/gemini-3-7-flash</a>
3.8 used 120M on high: <a href="https://artificialanalysis.ai/models/gemini-3-8-flash" rel="nofollow">https://artificialanalysis.ai/models/gemini-3-8-flash</a><p>Even their own chart showed more than 2x higher cost compared to 3.7: <a href="https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-3-8-cyber__evals__cwe-ben.width-2000.format-webp.webp" rel="nofollow">https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...</a></p>
]]></description><pubDate>Wed, 02 Sep 2026 17:07:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49539265</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49539265</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49539265</guid></item><item><title><![CDATA[New comment by eis in "Gemini 3.8 Flash and 3.8 Flash Cyber"]]></title><description><![CDATA[
<p>3.8 uses nearly twice as many tokens as 3.7. One might be inclined to think that they just uppsed the thinking budgets...<p>3.7 used 64M on high: <a href="https://artificialanalysis.ai/models/gemini-3-7-flash" rel="nofollow">https://artificialanalysis.ai/models/gemini-3-7-flash</a>
3.8 used 120M on high: <a href="https://artificialanalysis.ai/models/gemini-3-8-flash" rel="nofollow">https://artificialanalysis.ai/models/gemini-3-8-flash</a><p>Even their own chart showed more than 2x higher cost compared to 3.7: <a href="https://storage.googleapis.com/gweb-uniblog-publish-prod/images/gemini-3-8-cyber__evals__cwe-ben.width-2000.format-webp.webp" rel="nofollow">https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...</a></p>
]]></description><pubDate>Wed, 02 Sep 2026 17:06:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49539258</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49539258</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49539258</guid></item><item><title><![CDATA[New comment by eis in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>According to Artificial Analysis, 5.1 cost 56% <i>MORE</i> than 5, $8523 vs $5455. Yes cache cost is lower but it was <i>MUCH</i> more verbose: 140M vs 83M output tokens.<p>This directly contradicts what Anthropic is presenting here. Yes it scores higher but that's to be expected from a new release. It's the opposite of what OpenAI has been doing which was reducing costs, increasing efficiency.<p>Fable 5: <a href="https://artificialanalysis.ai/models/claude-fable-5" rel="nofollow">https://artificialanalysis.ai/models/claude-fable-5</a>
Fable 5.1: <a href="https://artificialanalysis.ai/models/claude-fable-5-1" rel="nofollow">https://artificialanalysis.ai/models/claude-fable-5-1</a></p>
]]></description><pubDate>Tue, 01 Sep 2026 20:50:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49528026</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49528026</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49528026</guid></item><item><title><![CDATA[New comment by eis in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>I am not sure if Fable is worth it, at least with version 5 vs Opus 5. Opus beats Fable in quite a few benchmarks and at twice the cost I just haven't seen it provide noticeably better results compared to Opus. Has anyone noticed big differences? I did notice Opus maybe making more mistakes repeatedly but I don't have hard numbers on this. I hope Fable 5.1 brings noticeable improvements. I am giving it a go now on my 20x Max plan on a problem that Opus 5 has struggled for more than week now and has made very slow progress with regular regressions on the way.</p>
]]></description><pubDate>Tue, 01 Sep 2026 18:11:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49525645</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49525645</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49525645</guid></item><item><title><![CDATA[New comment by eis in "I trained a small transformer in 1.5hrs and it beats many LLMs"]]></title><description><![CDATA[
<p>> Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning<p>Agreed and that's for any benchmark. Private tests are better but you still have to trust the provider to not log and use them for training.<p>That's why I like when a new set of tests like a new ARC-AGI version is published, that's where you can see which of the models abstracted to more general capabilities instead of being focused on the previous tasks. Most models completely fail new ARC-AGI tests.<p>The "67 cents" part though is misleading imho. You can't extrapolate from there and think that investing say $100 will get you a lot better results. You hit a ceiling very fast and investing into more compute will give you diminishing results. So yes, you can train a custom model to do somewhat decently on a specific set of tasks but then what?</p>
]]></description><pubDate>Tue, 01 Sep 2026 10:40:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49520167</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49520167</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49520167</guid></item><item><title><![CDATA[New comment by eis in "Konrad Zuse Museum shutting down due to lack of funding"]]></title><description><![CDATA[
<p>The person I replied to compared religion with physics (god vs deterministic computable universe). I said those are not remotely equally defensible theories.</p>
]]></description><pubDate>Mon, 31 Aug 2026 19:07:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49513646</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49513646</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49513646</guid></item><item><title><![CDATA[New comment by eis in "Konrad Zuse Museum shutting down due to lack of funding"]]></title><description><![CDATA[
<p>> Everything leads to paradoxes or unanswerable questions when you think it through. If our universe is computable then what "computer" is it running on? Why is our universe as computationally strong as it is and not more or less? Is it a simulation and if yes who is simulating it? Is there an infinite tower of universes each running progressively weaker (as in they can do less complex computation) versions of themselves? And if not then where does it stop and why?<p>None of the things you listed are paradoxes. A paradox is something that is self-contradictory.<p>> Even the deterministic part has problems, like why does quantum mechanics appear random when it's not or why we see ourselves as having free will.<p>Something can easily appear random when it is in fact not. Any random number your computer gives you is not truly random. Any hash looks random but is completely deterministic.<p>> A god is no better or worse than alien simulations, or an infinite multiverse, or the anthropic principle (aka giving up), or any other possible explanation for why the universe is the way it is.<p>Hard disagree. Not all attempts at explanations are equally valid or invalid. Some are more "out there" than others. And it should not stop us from trying to understand the universe more and more.</p>
]]></description><pubDate>Mon, 31 Aug 2026 19:05:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49513616</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49513616</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49513616</guid></item><item><title><![CDATA[New comment by eis in "Konrad Zuse Museum shutting down due to lack of funding"]]></title><description><![CDATA[
<p>The notion of god leads immediately to paradoxes and logical contradictions when thinking it through a little bit. It's fine if people have certain believes but let's not put religion on the same level as physics. No such paradoxes exist with the notion of a deterministic computable universe.<p>I concur with the OP when they said Zuse is being waaay underappreciated. He is one of the fathers of the modern computational age.</p>
]]></description><pubDate>Mon, 31 Aug 2026 17:19:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49512264</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49512264</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49512264</guid></item><item><title><![CDATA[New comment by eis in "NSA and IETF, Part 9"]]></title><description><![CDATA[
<p>I'm confused by your messages linked by DJB. You say that better cryptographers would not choose hybrids, which seems to say that you should indeed think that hybrids are not a good choice. Then you say you are not such a good cryptographer and would choose a hybrid. But if you know that more senior cryptographers think they are not the right choice then why choose them anyways? Or am I misreading "cryptography-literate" here?<p>Can you explain a bit more regarding your statement that DJB's POV on the matter has no broad support amongst his peers? I'm not in the field but Bernstein seemed like a highly respected member with a long track record in the crypto community, at least from the outside. Do you think the community is wrong or is it DJB who's wrong and why? There's also a good chance that I totally missed the argument being made.</p>
]]></description><pubDate>Sat, 15 Aug 2026 03:57:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49307503</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49307503</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49307503</guid></item><item><title><![CDATA[New comment by eis in "Gemini 3.7 Flash"]]></title><description><![CDATA[
<p>Sure, they are just checkpoints, that much I guess is obvious. The question is why did they not do frequent releases like this before and why are they making significant jumps in benchmarks so fast and all these companies suddenly falling into that pattern? Earning reports are not to come until end of October, that's not it.</p>
]]></description><pubDate>Thu, 13 Aug 2026 18:09:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49289841</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49289841</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49289841</guid></item><item><title><![CDATA[New comment by eis in "Gemini 3.7 Flash"]]></title><description><![CDATA[
<p>3.5 Pro was supposed to be around the corner two months ago. 4.0 Pro is some ways out as they recently stated they are seeing some promising early results from training. It didn't sound like a release is imminent.</p>
]]></description><pubDate>Thu, 13 Aug 2026 18:05:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49289769</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49289769</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49289769</guid></item><item><title><![CDATA[New comment by eis in "Gemini 3.7 Flash"]]></title><description><![CDATA[
<p>Grok, Meta, Gemini and others all released updates to their models within around a month or two from their respective last release and made significant jumps in benchmarks all around the same time. Any guesses as to why that is? Is it just the release season and/or everyone is benchmaxxing?</p>
]]></description><pubDate>Thu, 13 Aug 2026 17:46:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49289487</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49289487</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49289487</guid></item><item><title><![CDATA[New comment by eis in "When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation"]]></title><description><![CDATA[
<p>You post your benchmark on every other AI article, I've seen you do this by now more than a dozen times. It's a bit much. I don't want to be too harsh but your benchmark is obviously flawed when the top 3 models for Typescript (Combined) are Grok 4.5, Muse Spark 1.1 (lol), Gemini 3.5! Flash and then followed by Luna, beating Opus 5, Fable, 5.6 Sol etc by quite some margin. In fact 5.6 Sol ranks lower than Kimi K2.7 Code and even Grok Build 0.1. There are so many entries in your rankings that don't make any sense whatsoever that I can't take this benchmark serious and I have not seen it gaining traction. Please stop spamming it?</p>
]]></description><pubDate>Tue, 04 Aug 2026 21:07:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49175170</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=49175170</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49175170</guid></item><item><title><![CDATA[New comment by eis in "Grok 4.5"]]></title><description><![CDATA[
<p>Google wanted to release 3.5 Pro last month but because of the trouble Anthropic got with Fable they might have wanted to wait a bit for the dust to settle I could imagine. And now there is quite some competition. 3.5 Flash for me is a replacement to 3.1 Pro. It's more like a 3.2 Pro. It costs about the same (or more!) than 3.1 Pro, is a little bit smarter in many cases and a little bit faster.
3.5 Pro will be a lot more expensive and I expect it to juuuust be able to hang with Opus 4.8 and GPT-5.5.<p>I wish Google was able to actually push the industry further, either in terms of quality (intelligence) or quantity (price) but they've been playing catch up a lot.<p>They are playing the game a bit differently than all the others. The others have useable IDEs etc. while Google has a boatload of half-assed products.<p>Google better come out with a banger 3.5 Pro because who would have thought that Grok and GLM would be beating them?</p>
]]></description><pubDate>Wed, 08 Jul 2026 20:38:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48837131</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=48837131</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48837131</guid></item><item><title><![CDATA[New comment by eis in "Cloudflare Meerkat - Globally distributed consensus"]]></title><description><![CDATA[
<p>I feel like it would be much better if the article focused on QuePaxa because IMHO it's an algorithm that finally brings some novel ideas to concensus (e.g. not relying on timeouts) by kinda coming at it from a gossip protocol angle and is not getting the attention it deserves. The post shouldn't have focused and introduced Meerkat which hasn't been fully developed and tried in production. If they clearly presented the pros and cons vs not just Raft (which is popular but doesn't even play in the same league because it is relies on a leader) but other leaderless or multi-leader concensus protocols that would have been of greater value. The Paxos family of algorithms are a much closer fit here and there's a reason why some serious large planet scale systems choose it over Raft.<p>E.g. 1. Intro about issues with concensus 2. Intro to QuePaxa 3. Comparison to other algos that are close to it 4. Mentioning active work on implementation via Meerkat and intent to bring to production with followup posts.<p>As always when it comes to concensus it's all about trade-offs. And with QuePaxa that might be the increase in messages (note: I don't mean message round-trips). We'll see how it goes but it will definitely be interesting.</p>
]]></description><pubDate>Wed, 08 Jul 2026 16:06:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48833711</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=48833711</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48833711</guid></item><item><title><![CDATA[New comment by eis in "Cloudflare Meerkat - Globally distributed consensus"]]></title><description><![CDATA[
<p>But it's not a large-scale public deployment yet either. The article says towards the end that they just ran a proof of concept.<p>Maybe the blog post is just premature. It would be much more valuable if they posted it after actually having run it in production and validated the strengths and weaknesses with real world data.</p>
]]></description><pubDate>Wed, 08 Jul 2026 15:54:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48833573</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=48833573</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48833573</guid></item><item><title><![CDATA[Beijing is looking at curbing overseas access to China's top AI models]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/">https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48816025">https://news.ycombinator.com/item?id=48816025</a></p>
<p>Points: 64</p>
<p># Comments: 11</p>
]]></description><pubDate>Tue, 07 Jul 2026 10:51:07 +0000</pubDate><link>https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/</link><dc:creator>eis</dc:creator><comments>https://news.ycombinator.com/item?id=48816025</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48816025</guid></item></channel></rss>