<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Zababa</title><link>https://news.ycombinator.com/user?id=Zababa</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 28 Jul 2026 00:38:29 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Zababa" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by Zababa in "Apple Will 'Watch Everything Burn' When the AI Bubble Bursts"]]></title><description><![CDATA[
<p>>At their very core, Large Language Models' costs run contrary to basically every model of selling software.<p>>Consumers and enterprises alike have been trained to pay a monthly fee for a service, and while these services might have limits or strictures, basically nobody buying software expects to have a metered service, let alone one that's both metered and with hard to measure costs.<p>Has Ed Zitron not heard about the cloud? Unpredictable AWS bills?</p>
]]></description><pubDate>Mon, 27 Jul 2026 15:37:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49071157</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=49071157</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49071157</guid></item><item><title><![CDATA[New comment by Zababa in "The new rules of context engineering for Claude 5 generation models"]]></title><description><![CDATA[
<p>Interesting how everyone's favorite language seems to be even better in LLM era, almost like passion, skill level and having LLMs matters more than the language.</p>
]]></description><pubDate>Sun, 26 Jul 2026 11:44:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49057116</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=49057116</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49057116</guid></item><item><title><![CDATA[New comment by Zababa in "The new rules of context engineering for Claude 5 generation models"]]></title><description><![CDATA[
<p>Maybe this one will even be successful!</p>
]]></description><pubDate>Sat, 25 Jul 2026 22:10:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49052155</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=49052155</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49052155</guid></item><item><title><![CDATA[New comment by Zababa in "ARC-AGI Leaderboard"]]></title><description><![CDATA[
<p>>That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning.<p>They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.</p>
]]></description><pubDate>Sat, 25 Jul 2026 16:29:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49048978</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=49048978</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49048978</guid></item><item><title><![CDATA[New comment by Zababa in "Be skeptical of OpenAI's rogue hacker agent story"]]></title><description><![CDATA[
<p>>The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.<p>DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper.<p>You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff.</p>
]]></description><pubDate>Fri, 24 Jul 2026 20:37:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49041284</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=49041284</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49041284</guid></item><item><title><![CDATA[New comment by Zababa in "Does creatine make you smarter?"]]></title><description><![CDATA[
<p>I think the way you interpret the null result is downstream of considering creatine as "a supplement". You can make the null result say anything by changing your prior about creatine or supplements. That's the issue with priors.<p>You also seem to reject the possibility that some things help just a bit. The author has another article in the same vein about "things that can help maybe a bit but the evidence we have doesn't really help detecting small effects" <a href="https://dynomight.net/vitamin-d/" rel="nofollow">https://dynomight.net/vitamin-d/</a></p>
]]></description><pubDate>Wed, 22 Jul 2026 19:37:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49012266</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=49012266</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49012266</guid></item><item><title><![CDATA[New comment by Zababa in "Shinjuku Station in 3D"]]></title><description><![CDATA[
<p>Train station: :D :D :D <3<p>Train station, Japan: :D :D :D <3<p>Transit infrastructure is really cool</p>
]]></description><pubDate>Tue, 21 Jul 2026 12:34:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48991522</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48991522</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48991522</guid></item><item><title><![CDATA[New comment by Zababa in "Annoying and alarming things about OpenCode"]]></title><description><![CDATA[
<p>The mix of humanizing the LLM, calling it "clanker" and being very aggressive towards it is really weird. I don't think it's a good habit to take, it feels like it could bleed into how you interact with people. Many interactions are through text interfaces these days.</p>
]]></description><pubDate>Mon, 20 Jul 2026 14:15:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48979189</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48979189</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48979189</guid></item><item><title><![CDATA[New comment by Zababa in "Moonshine: Lets you stream games from your PC to any device running Moonlight"]]></title><description><![CDATA[
<p>There's also an alternative which is that token prices are already high, downtimes do exist (semi frequent on Claude) or are managed by serving degraded versions/quants (speculation that has never really be proven afaik) or reducing thinking time ("juice" values from open AI), and that free tiers are not used much at all or are a loss leader.</p>
]]></description><pubDate>Mon, 20 Jul 2026 13:07:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48978359</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48978359</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48978359</guid></item><item><title><![CDATA[New comment by Zababa in "What's the deal with all the random weekly quota resets for agents lately?"]]></title><description><![CDATA[
<p>>Resets of the weekly quota for all users must be ludicrously expensive for these companies<p>Why can't people see the alternative hypothesis, inference has huuuuuge margins?</p>
]]></description><pubDate>Sun, 19 Jul 2026 08:04:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=48965909</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48965909</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48965909</guid></item><item><title><![CDATA[New comment by Zababa in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>Interesting, good find! Yeah I may be wrong and this may be an error in the leaderboard. Weirdly it shows no reasoning cost and no reasoning tokens used, but for example here <a href="https://huggingface.co/datasets/arcprize/arc_agi_v1_public_eval/blob/main/deepseek-v3.2/1a2e2828.json" rel="nofollow">https://huggingface.co/datasets/arcprize/arc_agi_v1_public_e...</a> the answer is super short but it says "4945" completion tokens.</p>
]]></description><pubDate>Sat, 18 Jul 2026 12:12:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=48957389</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48957389</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48957389</guid></item><item><title><![CDATA[New comment by Zababa in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>I'm not referring to this paper, I'm referring to this leaderboard: <a href="https://arcprize.org/leaderboard">https://arcprize.org/leaderboard</a>. Set it to "arc agi 1", "base LLM" and you'll see deepseek at 57%. Submitted 2025-12-01, $0.120 per task. The paper you linked was later than that, and also says "We do not report an official ARC Prize leaderboard score".<p>So this paper doubled the price to get the same exact result at base Deepseek 3.2 at launch, and wasn't even tested on the verified set.</p>
]]></description><pubDate>Fri, 17 Jul 2026 18:18:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48950482</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48950482</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48950482</guid></item><item><title><![CDATA[New comment by Zababa in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.</p>
]]></description><pubDate>Fri, 17 Jul 2026 15:17:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48948431</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48948431</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48948431</guid></item><item><title><![CDATA[New comment by Zababa in "Mozilla: The state of open source AI"]]></title><description><![CDATA[
<p>Haven't tried Kimi K3 for now but there was a huge difference between GPT 5.6/Fable and GLM 5.2/Kimi K2.7 that were previous frontier open models.</p>
]]></description><pubDate>Fri, 17 Jul 2026 15:10:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48948342</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48948342</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48948342</guid></item><item><title><![CDATA[New comment by Zababa in "The state of open source AI"]]></title><description><![CDATA[
<p>The thing about not much difference between models and the harness making them deterministic and useful is wrong. Also models have different strengths and weaknesses and some are better at almost everything by a large margin compared to others.<p>As for your speculation, I think it's hinging on some companies releasing models for free or no big differences between models. In a world with hyperscalers and companies training models you can quickly recreate Anthropic or OpenAI by having an hyperscaler ally with a model training company, train a good/a better model, and not release it.</p>
]]></description><pubDate>Fri, 17 Jul 2026 15:09:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48948323</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48948323</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48948323</guid></item><item><title><![CDATA[New comment by Zababa in "Pebble Mega Update – July 2026"]]></title><description><![CDATA[
<p>This is true, but there's a big difference between saying "15-20 hours battery life, which is 11000 5-seconds activations, which last you a few years with 10 5-seconds activation a day" and "years of battery (btw in small text the real number is given). Especially since they mention that this project is hackable/you can do other things with it, knowing in advance you have something like ~100k button presses means some projects feel perfectly and some others won't really work.</p>
]]></description><pubDate>Fri, 17 Jul 2026 12:44:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48946715</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48946715</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48946715</guid></item><item><title><![CDATA[New comment by Zababa in "Pebble Mega Update – July 2026"]]></title><description><![CDATA[
<p>I don't really care about the environmental consciousness, my issue is that presenting a product with a battery that lasts for years when it actually lasts 15 to 20 hours makes me feel like I'm being lied to.</p>
]]></description><pubDate>Fri, 17 Jul 2026 08:21:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=48944666</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48944666</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48944666</guid></item><item><title><![CDATA[New comment by Zababa in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>"don't reinvent the wheel" isn't a law of physics and I think is mostly said by people that never designed anything with wheels.</p>
]]></description><pubDate>Fri, 17 Jul 2026 08:18:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48944640</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48944640</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48944640</guid></item><item><title><![CDATA[New comment by Zababa in "Schema Harness Achieves ~99% on Arc‑AGI‑3 Public"]]></title><description><![CDATA[
<p>I think I'd typify it as "ARC-AGI doesn't matter" more than "harness matters". Or maybe "harness matters for some very specific tasks".</p>
]]></description><pubDate>Fri, 17 Jul 2026 08:16:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48944622</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48944622</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48944622</guid></item><item><title><![CDATA[New comment by Zababa in "Pebble Mega Update – July 2026"]]></title><description><![CDATA[
<p>"Battery that lasts for years" being actually 12-15 hours of recording is a huge turn off honestly.<p>>How long does the battery last?<p>>Roughly 12 to 15 hours of recording. On average, I use it 10-20 times per day to record 3-6 second thoughts. That's up to 2 years of usage.<p>They then say:<p>>Wait, it's single use?<p>>Yes. We know this sounds a bit odd, but in this particular circumstance we believe it's the best solution to the given set of constraints. Other smart rings like Oura cost $250+ and need to be charged every few days. We didn't want to build a device like that. Before the battery runs out, the Pebble app notifies and asks if you'd like to order another ring.<p>My oura has lasted ~3 years, I recharge it twice a week usually, and I think it has spent way more than 15-20 hours turned on.</p>
]]></description><pubDate>Fri, 17 Jul 2026 08:06:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=48944549</link><dc:creator>Zababa</dc:creator><comments>https://news.ycombinator.com/item?id=48944549</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48944549</guid></item></channel></rss>