<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: jampa</title><link>https://news.ycombinator.com/user?id=jampa</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 11 Oct 2026 12:03:28 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=jampa" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by jampa in "The Test"]]></title><description><![CDATA[
<p>I think the reality is even simpler: Huang, Altman and Amodei are in the end of the day head salespeople of their respective company… and they use these social platforms to sell their products.<p>Their job is to convince people that AI has a large upside while minimizing any downside, they dont need to believe in those themselves.</p>
]]></description><pubDate>Fri, 25 Sep 2026 15:43:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49846150</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49846150</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49846150</guid></item><item><title><![CDATA[New comment by jampa in "Portal by Spotify cut my Claude Code token usage by 90%"]]></title><description><![CDATA[
<p>> I've never had an issue with Codex or Claude reading massive files<p>Reading files isn't a problem they want to solve. The idea seems to be using a cheaper model to "scout" for the intended code, instead of an expensive one that reads all the things (and spends more tokens / thinks about them).<p>I think this might be useful because Opus 5 especially tends to over-read. So this looks like an "LLM Bloom filter", telling "hey this is the code you might want to read".</p>
]]></description><pubDate>Sat, 05 Sep 2026 02:14:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49572440</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49572440</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49572440</guid></item><item><title><![CDATA[New comment by jampa in "Gemini 3.8 Flash and 3.8 Flash Cyber"]]></title><description><![CDATA[
<p>I wasn't trying to be precise originally, I just tried to fit activities into "morning / evening" buckets. I did the whole itinerary with Opus first, but when I gave it to Gemini 3.7 Flash to review, it started correcting it with "this place will close 5PM" or "this place is closed for good".<p>It was right on every nit, so it was surprising how well the model knows these things. If I ever release this I'll probably need the SERP API or Google Maps SDK (which I've heard is very expensive now), but for a personal trip where I will verify manually, using the LLM is okay for now.</p>
]]></description><pubDate>Wed, 02 Sep 2026 16:38:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49538832</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49538832</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49538832</guid></item><item><title><![CDATA[New comment by jampa in "Gemini 3.8 Flash and 3.8 Flash Cyber"]]></title><description><![CDATA[
<p>Eh that one is on me, if I think too much about my HN comment I end up deleting before posting it. I rely on the 1 min `delay` set in the profile page to fix before it goes live, but for some reason this time it was set to 0.</p>
]]></description><pubDate>Wed, 02 Sep 2026 16:30:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49538711</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49538711</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49538711</guid></item><item><title><![CDATA[New comment by jampa in "Gemini 3.8 Flash and 3.8 Flash Cyber"]]></title><description><![CDATA[
<p>I asked Claude to fix the grammar of my comment, and it changed "I am using 3.7 for" to "I've been using Claude 3.7", so they sneaked their own name on it.</p>
]]></description><pubDate>Wed, 02 Sep 2026 16:19:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49538550</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49538550</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49538550</guid></item><item><title><![CDATA[New comment by jampa in "Gemini 3.8 Flash and 3.8 Flash Cyber"]]></title><description><![CDATA[
<p>I've been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:<p>- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It's also the best at taking a cluster of places and working out a visiting order.<p>- Photo ranking (which photo should be the hero). Gemini can tell whether a photo is of the thing or of the view from it.<p>- Document parsing (extracting the relevant trip info from PDFs).<p>If you use LLMs for anything other than coding, I definitely recommend not discounting Gemini like I did just because other models are more popular.</p>
]]></description><pubDate>Wed, 02 Sep 2026 16:16:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49538512</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49538512</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49538512</guid></item><item><title><![CDATA[New comment by jampa in "Ask HN: What is one simple thing LLMs are insanely bad at?"]]></title><description><![CDATA[
<p>I tried doing something like this Ox Alpha with Opus advisor, having it work layer by layer (specs -> rooms -> room graphs ...), but each deliverable ended up a mess.<p>The curious thing is when I pointed out the flaws it fixed them quickly, but it's not something it can do without supervision, and supervising it takes more effort than doing the blueprint myself (to be fair, I'm not an architect, so I'm not the best at steering an LLM for this task).</p>
]]></description><pubDate>Wed, 26 Aug 2026 06:40:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49444898</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49444898</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49444898</guid></item><item><title><![CDATA[New comment by jampa in "Ask HN: What is one simple thing LLMs are insanely bad at?"]]></title><description><![CDATA[
<p>Serious answer: no model ever gets close to writing an architectural floor plan that makes sense.<p>They understand all the rules and best practices, they can (sometimes) spot a bad idea in a floor plan, they can describe a good floor plan.<p>But ask them to make one, even if you give it every detail (even a "node graph" of rooms), they will still output nonsense. Same for text and image models.<p>Floor plans should be the new Pelican Benchmark.</p>
]]></description><pubDate>Wed, 26 Aug 2026 05:01:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49444180</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49444180</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49444180</guid></item><item><title><![CDATA[New comment by jampa in "Opus 5.0 drives incoherence into the stratosphere"]]></title><description><![CDATA[
<p>Opus 5 feels like a downgrade from Opus 4.8 overall. It, along with Fable, really has a problem following instructions and staying in scope, and their prose keeps growing, both in explaining what it did and in writing multiline code comments (some comments read like a changelog, e.g. `// sky is blue (changed from red on 2026-01-01 per TCK-234 by @Foo)`).<p>Every time I ask it to do something, it does 80% of the job, goes off on "side quests" beyond the scope, and then leaves something out of the core ask (and when you tell it to finish, it does the same thing again).<p>The only advantage of Opus 5 over 4.8 is the better cutoff date for working with 3rd-party tools, though both do a very bad job of "this tool is constantly updated, I should look for the latest version first".</p>
]]></description><pubDate>Wed, 19 Aug 2026 18:22:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49365218</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49365218</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49365218</guid></item><item><title><![CDATA[New comment by jampa in "AI;DR (AI; Didn't Read)"]]></title><description><![CDATA[
<p>I used to do that, but the person could be bad at prompting too (which is often why the LLM couldn't give a good answer in the first place).<p>So the polite version I use now is: "Hey, just to get a bit more context, what was the original problem you were trying to solve?".<p>That gets them to distill their own problem a bit further.</p>
]]></description><pubDate>Mon, 17 Aug 2026 21:16:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49337761</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49337761</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49337761</guid></item><item><title><![CDATA[New comment by jampa in "Plug-in solar is coming. Plug-in batteries should follow"]]></title><description><![CDATA[
<p>The new rule here requires you to pay for part of your imports (up to ~60%), even if you export more. I know people who installed batteries in the inverters because of this. It is more expensive overall (hybrid inverter + LFP batteries) but has other benefits (blackouts are not a problem, most months you don't even use the grid). I imagine that once sodium-ion battery production scales and costs decline, this will become the "default" option for a solar installation.</p>
]]></description><pubDate>Sun, 02 Aug 2026 04:04:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49140992</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49140992</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49140992</guid></item><item><title><![CDATA[New comment by jampa in "Kimi K3-256k"]]></title><description><![CDATA[
<p>Anthropic has better SLA in Germany. I’ve heard uptime there can get up to nein nines.</p>
]]></description><pubDate>Wed, 29 Jul 2026 21:51:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49103548</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=49103548</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49103548</guid></item><item><title><![CDATA[New comment by jampa in "SpaceX wants to launch 100k more Starlink satellites for 100x the bandwidth"]]></title><description><![CDATA[
<p>When COVID hit, I knew a lot of engineers who decided to move to rural areas / small farms because they could leverage Starlink to work remotely.<p>Last year, when I asked whether they still liked Starlink, all of them said it is amazing, but they had gotten fiber coverage in their area from a local provider, so they don't use it anymore, or just use it as a backup.<p>I think Starlink was a huge demand signal that there were people willing to pay a premium for faster-than-radio internet. So, unless they manage to be cheaper and faster than fiber, I don't think there is much of an endgame there.<p>But there are a few places that will need Starlink, like planes, cruise ships, and islands. I'm just not sure if that will justify that $1T valuation.</p>
]]></description><pubDate>Fri, 10 Jul 2026 18:30:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=48863504</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48863504</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48863504</guid></item><item><title><![CDATA[New comment by jampa in "I think I have LLM burnout"]]></title><description><![CDATA[
<p>I feel the same way about consumer AI tools now. Gemini and ChatGPT have been abysmal lately. They can no longer be relied on to do multi-turn searching and thinking.<p>Before, they could stay in thinking mode for more than 7 minutes. For example, "find a source for this claim" would search, analyze, and self-adjust the query. Nowadays, even if I push for it, I cannot make these tools work for more than 30 seconds before they give generic answers, even in "Pro" mode.</p>
]]></description><pubDate>Thu, 09 Jul 2026 03:58:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48840813</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48840813</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48840813</guid></item><item><title><![CDATA[New comment by jampa in "Ask HN: Where is the programming profession going?"]]></title><description><![CDATA[
<p>From what you said: Not looking at code is bad, not because Claude can slip a few bugs (it can), but because LLMs tend to default to writing more code and features than needed, which isn't a good thing. I see a lot of people making 10+ PRs per day, but most of them are just going back to fix earlier PRs.<p>Claude always likes to "go big," for example, by choosing tools that can support millions of concurrent users or by adding unnecessary layers of abstraction that create more maintenance pain. I guess that's good for LLM companies, since more tokens are spent fixing the mess it caused.<p>Every time I enter plan mode for a huge feature, I end up cutting about 30-60% of the task scope before the LLM can actually start the work. I review the final code, and I still find things to cut. As said before "The best code is no code, or code you don’t have to maintain" [0]<p>0: <a href="https://www.simplethread.com/20-things-ive-learned-in-my-20-years-as-a-software-engineer/" rel="nofollow">https://www.simplethread.com/20-things-ive-learned-in-my-20-...</a></p>
]]></description><pubDate>Thu, 25 Jun 2026 04:12:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48668791</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48668791</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48668791</guid></item><item><title><![CDATA[New comment by jampa in "Why Is Claude Turning into an a**Hole?"]]></title><description><![CDATA[
<p>This post needs some examples, because I have never had an interaction with Claude that made me think this way.<p>LLMs generally have a way to "play a role" (most earlier prompt guides ask you to start with "You are a <role> expert in a <domain>"). So maybe if you interact with it by asking questions, it might assume that it knows more than the operator and adopt that attitude?</p>
]]></description><pubDate>Sun, 14 Jun 2026 22:31:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48533597</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48533597</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48533597</guid></item><item><title><![CDATA[New comment by jampa in "Claude Fable is relentlessly proactive"]]></title><description><![CDATA[
<p>I tested it to fix React Native bugs in a project, comparing it with Opus. It fared better on harder bugs, taking less time to find the root cause, but after implementing a fix, it spent a lot of time and effort on validation. This was mostly unnecessary, since most of the bugs were in the JS code, so for most things, hot reloading is enough for E2E validation and to run just the right tests. No need to run a full build and test suite (which takes 10+ minutes); the CI can do this.<p>I switched back to Opus because of this validation quirk. Overall, Fable spent 20% of the time on coding and 80% on validation.<p>I think using Fable for planning and Opus for execution could be a "best of both worlds" approach (I need to test this more), but for most cases, it's not necessary, and Opus is enough.</p>
]]></description><pubDate>Fri, 12 Jun 2026 02:35:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48499183</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48499183</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48499183</guid></item><item><title><![CDATA[New comment by jampa in "Claude Fable is relentlessly proactive"]]></title><description><![CDATA[
<p>Fable feels like a version of Opus running on a harness that won't let it halt until it's sure the issue is fixed, which makes sense if what you want is a model that's better at benchmarks.<p>It's a very good model, but it comes at a huge premium: not only do the tokens cost more, but the model itself really wants to spend them all. For example, working with React Native, Fable never just says "okay, I did the thing, that's it." It tries to rebuild the entire app from scratch, run the whole test suite, and watch every log and warning.<p>This is the first time with LLMs I've felt that upgrading to a model isn't worth it, even if my company lets me use it, because all the building / testing was just destroying my machine and its battery, which keeps me from working on other things.<p>For now, it feels like Opus with ultracode is a better choice (less pollution of the main context, more parallelism in investigations).</p>
]]></description><pubDate>Fri, 12 Jun 2026 02:01:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48498951</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48498951</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48498951</guid></item><item><title><![CDATA[New comment by jampa in "What it feels like to work with Mythos"]]></title><description><![CDATA[
<p>It is hallucinating many flights in my region, some that never existed (so it is not an outdated data problem).<p>I also see some logic flaws. It overlooks the option of going to a major hub to access faster aircraft, rather than hopping on local hubs.<p>Also, immigration and customs are cleared at the first airport you arrive at in the country, not at the last one.<p>In some countries, you need to clear immigration even while going to a third country, so 1 hour is not enough to do it.</p>
]]></description><pubDate>Tue, 09 Jun 2026 19:55:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48466758</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48466758</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48466758</guid></item><item><title><![CDATA[New comment by jampa in "Wind and solar generated more power than gas globally in April 2026"]]></title><description><![CDATA[
<p>> there's not enough money to be made via speculation<p>I mean, there is money to be made. CATL stock (the major producer of EV batteries with 50% market share, with billions of contracts for stationary batteries) rose 48.81% over the last 6 months, for example.<p>But I agree that news about renewables goes unnoticed. I only see news about renewables because I actively seek out channels and websites that cover it. I wonder if it is because most companies in the industry are Chinese and don't focus on PR in the West as AI companies do.</p>
]]></description><pubDate>Thu, 04 Jun 2026 15:24:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48400059</link><dc:creator>jampa</dc:creator><comments>https://news.ycombinator.com/item?id=48400059</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48400059</guid></item></channel></rss>