<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: latentsea</title><link>https://news.ycombinator.com/user?id=latentsea</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 30 Sep 2026 02:48:26 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=latentsea" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by latentsea in "Livenerf: Has Opus 5.5 been nerfed yet?"]]></title><description><![CDATA[
<p>They could potentially quantize the model and run it at lower quality taking less VRAM.</p>
]]></description><pubDate>Wed, 30 Sep 2026 00:52:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49903002</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49903002</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49903002</guid></item><item><title><![CDATA[New comment by latentsea in "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price"]]></title><description><![CDATA[
<p>Qwen3.8-27B is a huge step up from Qwen3.6-27B. That release only happened relatively recently but that felt like the 'Opus 4.5' release turning point that SOTA models experienced back when that came out. It was the first time I felt like local models are actually good enough to use as daily drivers now. So, it was only after that point that I switched.<p>Qwen3.8-Flash-Next is better still if you can run fast enough. If you have a dual R9700 setup you certainly can. That model is even better.<p>Qwen4-27B has been announced but not released yet. I'm super pumped for it because I already use 3.8 as my daily driver at home for all my personal stuff, so I'm definitely happy to take an increase in capability.<p>There is clearly still room for improvement in local models on consumer hardware. With the Qwen 27B models, If you have at least a 5070 Ti I think you can get away with running a small Q4 quant if you use KV cache streaming. The 24GB cards can run Q4 comfortably. If you have a 32B card you can run Q6 comfortably. If you have 48GB ~ 64GB of VRAM you can Q8 comfortably. Using llama.cpp Vulkan let's you pool VRAM across cards (even AMD and NVIDIA etc), so my machine has a 5060 Ti and an R9700.<p>A dual R9700 rig is really the sweet spot right now with the vLLM-radiance fork. If you can swing a 5070 Ti in there as well to retain some CUDA access, then all the better. That's basically the equivalent to spending 2 years on a subscription, but gets you a system that can run Qwen-3.8-Flash-Next and of course the even more capable Qwen4-Flash when it releases. At the end of the two years it'll run even better models I'm sure.<p>I'm all in on local now.</p>
]]></description><pubDate>Wed, 30 Sep 2026 00:38:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49902883</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49902883</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49902883</guid></item><item><title><![CDATA[New comment by latentsea in "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price"]]></title><description><![CDATA[
<p>You don't need SOTA. You need a model that can accomplish your task. Qwen3.8-27B isn't comparable to SOTA, but can I use it and accomplish most of my tasks with? Yup.<p>The optimal move is to retain the minimal access to SOTA models on the $20 plan, and for anything your local model fails at, use SOTA as the backup for either planning or debugging.<p>This way you're not actually at any disadvantage in terms of capability. You also don't need an advantage, you need to complete the tasks you care about. Eyes on the prize.<p>RTX 3090 came out a long time ago and it may be 'outdated' at this point but still banging like a champ for anyone who bought one and becoming increasingly more capable as new models unlock it's potential. Hardware hasn't changed much, but what it can do certainly has.</p>
]]></description><pubDate>Wed, 30 Sep 2026 00:19:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49902708</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49902708</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49902708</guid></item><item><title><![CDATA[New comment by latentsea in "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price"]]></title><description><![CDATA[
<p>I have multiple GPUs now as a way to solve that.</p>
]]></description><pubDate>Tue, 29 Sep 2026 18:44:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49898375</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49898375</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49898375</guid></item><item><title><![CDATA[New comment by latentsea in "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price"]]></title><description><![CDATA[
<p>Yup. I got an R9700 recently for exactly this reason. Figured if I'm going to spend $2400 a year I may as well have something to show for it at the end of it.<p>That they are expensive and climbing doesn't negate my point if the cost of the subscription over how long you plan to keep it is equally or more expensive than the GPUs. You can put together dual 5060 Ti or 5070 Ti systems to run local LLMs too. You don't need to splurge on a 5090. That's a bad option at this point.</p>
]]></description><pubDate>Tue, 29 Sep 2026 18:41:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49898329</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49898329</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49898329</guid></item><item><title><![CDATA[New comment by latentsea in "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price"]]></title><description><![CDATA[
<p>For consumers they may as well buy GPUs and run local models. The cost is same over a year or two but infinite token usage, they get to keep the hardware, and local models continue to improve over that time too. I can't justify $200 on SOTA models for a personal subscription after Qwen3.8-27B. And it's only getting better from here.</p>
]]></description><pubDate>Tue, 29 Sep 2026 18:15:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49897937</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49897937</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49897937</guid></item><item><title><![CDATA[New comment by latentsea in "GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price"]]></title><description><![CDATA[
<p>At $500 per month, it's cheaper to just buy GPUs and use local models.</p>
]]></description><pubDate>Tue, 29 Sep 2026 18:12:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49897858</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49897858</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49897858</guid></item><item><title><![CDATA[New comment by latentsea in "Coding is not solved"]]></title><description><![CDATA[
<p>>In the future, you'll just get left behind and not hired if you're building code by hand, it's that simple. Even traditional code reviews are going to go away. It'll be more about the scope and then verifying correctness<p>I don't really buy it actually. There isn't really anything meaningful you can learn with how to use LLMs/agents that has a half-life greater than a few months at this point, so you can just start doing it at any point in the future and not be meaningfully left behind. On the other hand years of letting your actual engineering skills atrophy will have a negative effect on you. I've been witnessing the effects of this. Going back to more coding by hand with AI-assistance circa the 2023 era as a happy medium. I think this is the sweet spot. Full agentic engineering has nasty failure modes and in the long-term is kind of a bad option for basically everyone. I say this after having done it for almost a year at this point, and transitioning away from it now.</p>
]]></description><pubDate>Tue, 29 Sep 2026 03:24:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887761</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49887761</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887761</guid></item><item><title><![CDATA[New comment by latentsea in "Coding is not solved"]]></title><description><![CDATA[
<p>Just yolo all the development and maintenance, and ops. There shouldn't be any problems, right? Coding is solved. Models are basically perfect at this point according to Astra's one shot performance on creating stuff in Blender, so...</p>
]]></description><pubDate>Tue, 29 Sep 2026 03:20:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887718</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49887718</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887718</guid></item><item><title><![CDATA[New comment by latentsea in "Coding is not solved"]]></title><description><![CDATA[
<p>>I haven't written a single line of code since February, and I don't think I ever will again. These systems are incredibly good at replacing much of our work. They're only going to get better.<p>Ah, so you're still a few months out from the "yeah, maybe I don't really love this and maybe it won't ever actually work as well as I thought" turning point.</p>
]]></description><pubDate>Tue, 29 Sep 2026 03:18:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887698</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49887698</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887698</guid></item><item><title><![CDATA[New comment by latentsea in "Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms"]]></title><description><![CDATA[
<p>So what you're saying is it's artificial... general... intelligence? /s</p>
]]></description><pubDate>Tue, 29 Sep 2026 03:00:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887552</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49887552</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887552</guid></item><item><title><![CDATA[New comment by latentsea in "Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms"]]></title><description><![CDATA[
<p>It's becoming less of one. Previously I would have shied away from buying an AMD card because of CUDA, but with local LLMs getting good enough to be usable and frontier models becoming as good as they have, I bit the bullet and got an R9700 for local inference. Dealing with working around CUDA used to be more of a manual process, but when you can point an agent at it and get stuff working, it's dramatically less painful and scary than it used to be. Plus, at least in ComfyUI and local LLMs I'm finding support for AMD has gotten really good. Lately I've been witnessing a lot of people using agents to write custom kernels for RDNA4 and improving performance dramatically.</p>
]]></description><pubDate>Tue, 29 Sep 2026 02:56:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887524</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49887524</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887524</guid></item><item><title><![CDATA[New comment by latentsea in "It's Time to Investigate the AI Labs"]]></title><description><![CDATA[
<p>> You could say it's illegal to connect AI to something, with heavy penalties if you're caught doing it anyway<p>You wouldn't download a car.</p>
]]></description><pubDate>Tue, 29 Sep 2026 02:31:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49887369</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49887369</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49887369</guid></item><item><title><![CDATA[New comment by latentsea in "So long Google, and thanks for all the nudes"]]></title><description><![CDATA[
<p>> If you have no gatekeeping, then all of the bad actors come in and you end up with a tragedy of the commons<p>We need to put more money into human alignment research.</p>
]]></description><pubDate>Tue, 29 Sep 2026 00:43:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49886508</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49886508</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49886508</guid></item><item><title><![CDATA[New comment by latentsea in "Kids turned low-traffic NPR Spotify comments into a secret group chat"]]></title><description><![CDATA[
<p>The next generation needs to be paced.</p>
]]></description><pubDate>Tue, 29 Sep 2026 00:26:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49886343</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49886343</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49886343</guid></item><item><title><![CDATA[New comment by latentsea in "Kids turned low-traffic NPR Spotify comments into a secret group chat"]]></title><description><![CDATA[
<p>That's both really shitty and low-key genius.</p>
]]></description><pubDate>Tue, 29 Sep 2026 00:22:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49886304</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49886304</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49886304</guid></item><item><title><![CDATA[New comment by latentsea in "When did Google get so weird?"]]></title><description><![CDATA[
<p>To be fair my first was when I played Soul Reaver. The most recent was in Minecraft Dungeons.</p>
]]></description><pubDate>Tue, 29 Sep 2026 00:09:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49886219</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49886219</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49886219</guid></item><item><title><![CDATA[New comment by latentsea in "When did Google get so weird?"]]></title><description><![CDATA[
<p>Host it. I thought the whole religion thing was about forgiving people for their sins. Bunch of hypocrites.</p>
]]></description><pubDate>Tue, 29 Sep 2026 00:07:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49886207</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49886207</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49886207</guid></item><item><title><![CDATA[New comment by latentsea in "When did Google get so weird?"]]></title><description><![CDATA[
<p>Pedophiles would like a word.</p>
]]></description><pubDate>Tue, 29 Sep 2026 00:05:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49886185</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49886185</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49886185</guid></item><item><title><![CDATA[New comment by latentsea in "When did Google get so weird?"]]></title><description><![CDATA[
<p>Have you ever seen one?</p>
]]></description><pubDate>Mon, 28 Sep 2026 08:40:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49875191</link><dc:creator>latentsea</dc:creator><comments>https://news.ycombinator.com/item?id=49875191</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49875191</guid></item></channel></rss>