<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: jug</title><link>https://news.ycombinator.com/user?id=jug</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 26 Aug 2026 13:12:39 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=jug" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by jug in "Ox Alpha"]]></title><description><![CDATA[
<p>Political? It's using a Chinese LLM trait. If it only refused to talk about a particular kind of soccer, we'd probe it with that instead. The goal is not to discuss politics, the goal is to find out which model it is. What's political here is the language model.</p>
]]></description><pubDate>Fri, 21 Aug 2026 11:05:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49386371</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49386371</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49386371</guid></item><item><title><![CDATA[New comment by jug in "Meta faces 'astronomical' consequences as legal fight reaches critical moment"]]></title><description><![CDATA[
<p>From what I'm seeing I do believe TikTok is worse.</p>
]]></description><pubDate>Wed, 19 Aug 2026 16:45:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49363936</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49363936</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49363936</guid></item><item><title><![CDATA[New comment by jug in "GPT 5.6 Sol is the best "vision" model OpenAI ever released"]]></title><description><![CDATA[
<p>I really like the combo 5.6 Luna & Sol for price and performance and would be perfectly happy if they stayed here for a moment without mucking about with sidegrades that I think AI evolution has often felt like lately.</p>
]]></description><pubDate>Mon, 17 Aug 2026 14:41:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49331837</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49331837</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49331837</guid></item><item><title><![CDATA[New comment by jug in "DeepSeek V4 Pro 0813"]]></title><description><![CDATA[
<p>This is true and is only becoming more important the more they improve. I am already moving to checking so they're at least somewhat following the status quo and otherwise prioritizing price and platform. I think this will be an emerging way of viewing AI in 2027 and the winner will probably be open models and China.</p>
]]></description><pubDate>Wed, 12 Aug 2026 22:58:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49279721</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49279721</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49279721</guid></item><item><title><![CDATA[New comment by jug in "DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis"]]></title><description><![CDATA[
<p>If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. cost efficient models aren't necessarily invalidated early by progress that matters.</p>
]]></description><pubDate>Sun, 02 Aug 2026 00:05:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49139824</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49139824</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49139824</guid></item><item><title><![CDATA[New comment by jug in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost efficient way, and especially now!</p>
]]></description><pubDate>Fri, 31 Jul 2026 00:51:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49117793</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49117793</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49117793</guid></item><item><title><![CDATA[New comment by jug in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>Kimi K3 is fairly cheap per token but thinks like a madman with poor self esteem.</p>
]]></description><pubDate>Fri, 31 Jul 2026 00:48:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49117776</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49117776</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49117776</guid></item><item><title><![CDATA[New comment by jug in "Enabling two settings tripled our scores on the ARC-AGI-3 benchmark"]]></title><description><![CDATA[
<p>So is this how Opus 5 ran it?<p><a href="https://arcprize.org/results/anthropic-claude-opus-5">https://arcprize.org/results/anthropic-claude-opus-5</a></p>
]]></description><pubDate>Thu, 30 Jul 2026 11:59:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49108831</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49108831</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49108831</guid></item><item><title><![CDATA[New comment by jug in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.</p>
]]></description><pubDate>Mon, 27 Jul 2026 13:05:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49069140</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49069140</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49069140</guid></item><item><title><![CDATA[New comment by jug in "Opus 5 is currently #1 on Artificial Analysis Intelligence Leaderboard"]]></title><description><![CDATA[
<p>I agree and this is why I think open models will win in the end. There is just so much to gain on being 10% behind the curve. Especially when the curve is far beyond your needs.</p>
]]></description><pubDate>Sun, 26 Jul 2026 19:56:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49061829</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49061829</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49061829</guid></item><item><title><![CDATA[New comment by jug in "Claude Opus 5"]]></title><description><![CDATA[
<p>This will also make it harder to compete because it holds true for everyone, not just you. I am already, this year, seeing vibe coders put out some decent stuff on App Stores but the problem is marketing it. You no longer automatically stand out just because you have a cool app. You may not even do so with reasonable early traction via social media. And then what? Throw venture capital onto the problem? In an AI saturated world? Really?</p>
]]></description><pubDate>Sat, 25 Jul 2026 10:41:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49046409</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=49046409</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49046409</guid></item><item><title><![CDATA[New comment by jug in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>Yeah I've noted this behavior with best in class open weight models. They said K3 would have token efficiency improvements and I was hoping especially solving the thinking loop issue that plagued K2.x but even if this release helped somewhat, it looks like we still have a long way to go here... I'm not sure what's up here but I suppose lacking finetuning quality.<p>What OpenAI in particular have done with reasoning efficiency in the past few months since ChatGPT 5.5 is nothing short of remarkable. It's overshadowed a bit by the benchmark game and the Fable hoopla.<p>Now is the time to focus less on token cost and intelligence, but tokens to solve a particular set of tasks in closed benchmarks for a variety of categories.<p>What is the use of grand intelligence if it either costs you a kidney or can't complete at all within a token budget? Even if there are niche uses where you truly want "maximum power" above all, we need to at least more severely penalize such models versus those that does it just as fine within a tenth of the token cost.<p>I'm aware of some benchmarks at the Artificial Intelligence site, but CLEARLY we are not focusing enough on these today and still leaving the fun surprises to the users.</p>
]]></description><pubDate>Sat, 18 Jul 2026 20:21:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48961942</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48961942</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48961942</guid></item><item><title><![CDATA[New comment by jug in "Kimi K3, and what we can still learn from the pelican benchmark"]]></title><description><![CDATA[
<p>I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the headline feature, the headline benchmark, etc etc. At least looking at Anthropic, Google, OpenAI. There are of course other LLM's.<p>So then add a dash of cybersecurity and medical use and that's basically it. No "closer to AGI" advertising. I'd say the 2026 development has in fact been the opposite; optimizing AI for niches where there is most potential for profits and that your description died in circa GPT-5 era.<p>In fact, this problem (for this test) is also stated by the pelican test author:<p>"The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.<p>So don’t go using pelicans to compare models!"</p>
]]></description><pubDate>Sat, 18 Jul 2026 10:24:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48956697</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48956697</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48956697</guid></item><item><title><![CDATA[New comment by jug in "I think I have LLM burnout"]]></title><description><![CDATA[
<p>Addiction due to the dopamine hits of occasional struggle and then churning out apps that work: <a href="https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-paying-the-price?utm_id=97760_v0_s00_e0_tv4" rel="nofollow">https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...</a></p>
]]></description><pubDate>Thu, 09 Jul 2026 15:31:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48847595</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48847595</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48847595</guid></item><item><title><![CDATA[New comment by jug in "I think I have LLM burnout"]]></title><description><![CDATA[
<p>This article is also related to exhausting AI through generating pressure and posted here recently:<p>AI coding is addictive. Engineers are paying the price
<a href="https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-paying-the-price" rel="nofollow">https://leaddev.com/ai/ai-coding-is-addictive-engineers-are-...</a></p>
]]></description><pubDate>Thu, 09 Jul 2026 15:22:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=48847470</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48847470</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48847470</guid></item><item><title><![CDATA[New comment by jug in "Grok 4.5"]]></title><description><![CDATA[
<p>You probably have it backwards. It's Grok that is shoving right wing ideology down your throat. Research has shown that without specific guidance to otherwise, LLM's tend to be slightly left leaning by default. There are some theories as for why this is so.</p>
]]></description><pubDate>Thu, 09 Jul 2026 00:25:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48839330</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48839330</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48839330</guid></item><item><title><![CDATA[New comment by jug in "GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday"]]></title><description><![CDATA[
<p>Yet, in a month we'll be fine. We were fine with Anthropic naming models by music. I'm sure celestial bodies will be OK too. Larger = better. It's simple. As for the why? Marketing, making products feel "fresh", exciting, new, something alluring that we didn't have before. So, much like since industrialization.<p>What surprises me is not this, but that OpenAI changed things up without syncing with a GPT 6.</p>
]]></description><pubDate>Wed, 08 Jul 2026 09:13:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=48829531</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48829531</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48829531</guid></item><item><title><![CDATA[New comment by jug in "Odin 1.0 Announcement"]]></title><description><![CDATA[
<p>Free tier of Google Gemini can summarize and let you ask questions about pasted YT links.</p>
]]></description><pubDate>Tue, 07 Jul 2026 14:08:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=48818110</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48818110</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48818110</guid></item><item><title><![CDATA[New comment by jug in "60% Fable cost cut by converting code to images and having the model OCR it"]]></title><description><![CDATA[
<p>Alternative 1 isn’t all that unlikely given Opus 4.8 couldn’t do this. So it’s a recently possible hack. Not something LLM corps have been blindsided by for years. I also strongly recommend RTFA in this case, namely ”The honest part, read before relying on it”</p>
]]></description><pubDate>Fri, 03 Jul 2026 18:35:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=48778314</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48778314</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48778314</guid></item><item><title><![CDATA[New comment by jug in "Kimi K2.7 Code is generally available in GitHub Copilot"]]></title><description><![CDATA[
<p>I often feel like we're nowadays mostly pushing AI developments in the ways of finetuning differences. Like how new editions of Claude are tuned for agentic coding which might even be detrimental if you're using it for non-agentic coding. Or how Fable 5 in fact do look great but at a huge cost for inference and a high likelihood of post-launch nerfs or limit/price revisions. How Gemini 3.5 has more liberal limits but on the other hand underperforms a bit.<p>It's like we're mostly treading mud at this point. New editions are released, a version number increases, but I have to wonder if all steps are forward or they're more just tuned differently with similar actual perf per dollar as when this year began.<p>Most in fact seem to be happening to me with small models. Like your Qwen. Or Gemma 4 31B which is kinda magic especially when considering multilingual abilities. So yes, in that sense I can see "development" probably as we refine data sets and training methods but I see it less on the big hulking beasts with daily limits (unless you turn it up to 11 like Fable).<p>Edit: As I posted this, I saw a "before and after" comparison for Fable and the reintroduced version is seeing a catastrophic drop in BridgeBench performance as they're still mucking with the model. Go figure... <a href="https://x.com/Hesamation/status/2072692225100612032" rel="nofollow">https://x.com/Hesamation/status/2072692225100612032</a></p>
]]></description><pubDate>Thu, 02 Jul 2026 18:14:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48765339</link><dc:creator>jug</dc:creator><comments>https://news.ycombinator.com/item?id=48765339</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48765339</guid></item></channel></rss>