<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: MrScruff</title><link>https://news.ycombinator.com/user?id=MrScruff</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Mon, 28 Sep 2026 09:02:52 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=MrScruff" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by MrScruff in "Fuck it, make it anyway"]]></title><description><![CDATA[
<p>The distinction that can be drawn is between the code, and the user experience of the software. You may argue that the two are inextricably linked, but I don’t think that can be treated as an absolute principle.<p>I'd also say I've worked with many engineers over the years who were more focused on the software engineering aspect that the user experience. Depending on the use case for the software, that can absolutely end up being the limiting factor.</p>
]]></description><pubDate>Sun, 13 Sep 2026 08:20:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49681389</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49681389</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49681389</guid></item><item><title><![CDATA[New comment by MrScruff in "Fuck it, make it anyway"]]></title><description><![CDATA[
<p>It's a bad analogy, because unless you're cooking a recipe, cooking is about all the small decisions you're making during the process which could apply both to coding or the design aspect of building software with an LLM.</p>
]]></description><pubDate>Sat, 12 Sep 2026 17:28:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49674764</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49674764</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49674764</guid></item><item><title><![CDATA[New comment by MrScruff in "Fuck it, make it anyway"]]></title><description><![CDATA[
<p>I don't agree. That would imply they asked for a piece of software and the LLM one-shotted it, but that's absolutely not what's happening in the vast majority of cases. If two people vibe code the same tool from the same high level desciption the results are going to be quite different.</p>
]]></description><pubDate>Sat, 12 Sep 2026 17:20:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49674669</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49674669</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49674669</guid></item><item><title><![CDATA[New comment by MrScruff in "OpenAI agents carried out an undisclosed attack on RubyGems"]]></title><description><![CDATA[
<p>I think we’re seeing the agents become very advanced at tasks with verifiable reward through RL. Currently they don’t exhibit the same skills in their attempts to manipulate humans - presumably because they’re not being specifically trained for that. But they are certainly not aligned in the sense that they will attempt social engineering, they’re just not very good at it (yet).<p>However, if in the future AIs become much more efficient at learning without requiring vast amounts of RL, closer to how humans learn. Then you would have to assume we’d have a real problem.</p>
]]></description><pubDate>Sat, 12 Sep 2026 09:31:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49670624</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49670624</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49670624</guid></item><item><title><![CDATA[New comment by MrScruff in "An Alien Mind"]]></title><description><![CDATA[
<p>The point is, the big improvements we’re seeing nowadays are coming from RL, not from scraping the internet.</p>
]]></description><pubDate>Sun, 06 Sep 2026 20:58:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49590875</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49590875</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49590875</guid></item><item><title><![CDATA[New comment by MrScruff in "AI, Tools and Transformation"]]></title><description><![CDATA[
<p>I think this all rings true for where we are right now. The trend is that the agents are becoming superhuman in tasks for which there is a verifiable reward, and analysing a business problem, identifying inefficiencies and turning it into a software specification is not one of them.<p>However, things are changing so rapidly that I can see that starting to change as well. But it would take much better learning efficiency to understand unknown domains, 100% computer use reliability etc. I suspect we’ll see this by the end of the decade.</p>
]]></description><pubDate>Sun, 06 Sep 2026 09:45:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49584824</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49584824</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49584824</guid></item><item><title><![CDATA[New comment by MrScruff in "“Next-token predictor” is the wrong mental model for LLMs"]]></title><description><![CDATA[
<p>The point was, if your internal model of the world makes a prediction of a negative outcome at some point in the future, and you optimise your individual actions to avoid that negative outcome, then wouldn’t it make sense to focus on the fact you’re building and optimizing towards an internal world model rather than the fact you’re executing your actions one at a time in series?</p>
]]></description><pubDate>Sat, 05 Sep 2026 08:41:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49574532</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49574532</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49574532</guid></item><item><title><![CDATA[New comment by MrScruff in "“Next-token predictor” is the wrong mental model for LLMs"]]></title><description><![CDATA[
<p>I am not an expert, but I do understand the distinction that is being made here. It makes sense to describe the result of pre-training as a ‘next token’ predictor as that’s what it’s been trained to do, not because it’s an autoregressive architecture that produces tokens one at a time.<p>If this base is then trained using RL towards a different objective (maths and coding), the model becomes fundamentally a different thing and the recent models are clear evidence of that, regardless of they fact they remain autoregressive.</p>
]]></description><pubDate>Sat, 05 Sep 2026 08:18:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49574337</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49574337</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49574337</guid></item><item><title><![CDATA[New comment by MrScruff in "“Next-token predictor” is the wrong mental model for LLMs"]]></title><description><![CDATA[
<p>Not sure if this was a serious comment but it’s worth considering that humans have a long history of figuring out ways to make other humans work for them without bestowing rights on them.</p>
]]></description><pubDate>Sat, 05 Sep 2026 08:06:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49574255</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49574255</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49574255</guid></item><item><title><![CDATA[New comment by MrScruff in "AI Can Make You Suck Faster Too"]]></title><description><![CDATA[
<p>In general, the frontier models are not capable of reliably authoring non-trivial code without careful oversight yet. They are great at producing code that can pass tests, but not neccessarily a code review. This means if you care about code quality you still need a human in a loop understanding what has been done, and that becomes the bottleneck. And less disciplined folks will indeed become increasingly dependent.<p>However, over time the complexity of problems where you can get away with less/no oversight is increasing. And the models are already great at solving certain classes of problems where one doesn't really care that much about code quality, that wouldn't have even been attempted in a pre-LLM world. Over the weekend I was using Claude to add features to the compiled (no source available) firmware of one of my audio devices, adding workflow features by patching assembly and custom DSP code.<p>In coding, as with other areas, what's emerging is jagged intelligence.</p>
]]></description><pubDate>Tue, 01 Sep 2026 07:08:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49518923</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49518923</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49518923</guid></item><item><title><![CDATA[New comment by MrScruff in "Why your local LLM feels dumber than it is"]]></title><description><![CDATA[
<p>I get around 20 tok/s, 4 bit quant, MTP, 4 bit KV cache quantisation. On an M4 Pro 48Gb.</p>
]]></description><pubDate>Sat, 22 Aug 2026 21:46:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49404188</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49404188</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49404188</guid></item><item><title><![CDATA[New comment by MrScruff in "llama.cpp"]]></title><description><![CDATA[
<p>I thought the main advantage of oMLX is it's less likely to invalidate the KV cache when working with coding agents, which is key when working on a Mac because of the slower prompt processing.</p>
]]></description><pubDate>Wed, 12 Aug 2026 07:58:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49269148</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49269148</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49269148</guid></item><item><title><![CDATA[New comment by MrScruff in "H3-metal – Native MiniMax-H3 inference for Apple Silicon"]]></title><description><![CDATA[
<p>LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.</p>
]]></description><pubDate>Tue, 11 Aug 2026 12:45:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49257410</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49257410</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49257410</guid></item><item><title><![CDATA[New comment by MrScruff in "LLMs reward expertise"]]></title><description><![CDATA[
<p>I got my girlfriend to install Claude Code and she was happily able to create software with it completely independently of me.</p>
]]></description><pubDate>Tue, 04 Aug 2026 07:55:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49165564</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=49165564</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49165564</guid></item><item><title><![CDATA[New comment by MrScruff in "Maybe you should learn something"]]></title><description><![CDATA[
<p>Yeah exactly. After a hard day when my brain is frazzled, a workout will actually make me feel better.</p>
]]></description><pubDate>Sat, 04 Jul 2026 15:12:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=48785983</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=48785983</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48785983</guid></item><item><title><![CDATA[New comment by MrScruff in "Maybe you should learn something"]]></title><description><![CDATA[
<p>I think what the parent post was saying is that there is a finite amount of useful mental function time in any one day, and once you’ve exhausted this any attempted learning will be pretty inefficient. Also some jobs will have a faster burn rate. Doing a workout is separate as it doesn’t draw on the mental energy pool.</p>
]]></description><pubDate>Sat, 04 Jul 2026 14:11:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=48785546</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=48785546</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48785546</guid></item><item><title><![CDATA[New comment by MrScruff in "AI is 'not smart' so what's next in artificial intelligence?"]]></title><description><![CDATA[
<p>Considering all of the great research that has come from his labs (eg. DINO, Segment Anything) I don’t think that’s fair (no pun intended).</p>
]]></description><pubDate>Fri, 03 Jul 2026 09:04:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48772675</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=48772675</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48772675</guid></item><item><title><![CDATA[New comment by MrScruff in "Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?"]]></title><description><![CDATA[
<p>I’m normally comparing frontier open/cheap models against frontier closed source. I use deepseek/glm regularly, they’re fine and you can get real work done with them but it’s super obvious when you switch back to opus or even sonnet. A 3B active param MoE model is not comparable.</p>
]]></description><pubDate>Mon, 15 Jun 2026 19:57:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=48546256</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=48546256</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48546256</guid></item><item><title><![CDATA[New comment by MrScruff in "Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?"]]></title><description><![CDATA[
<p>You really need to take the benchmarks with a massive pinch of salt. I’ve been testing local LLMs since the original llama and there’s nothing I’ve tried that is in the same category as Opus.</p>
]]></description><pubDate>Mon, 15 Jun 2026 19:41:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48546055</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=48546055</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48546055</guid></item><item><title><![CDATA[New comment by MrScruff in "Shepherd's Dog: A Game by the Most Dangerous AI Model"]]></title><description><![CDATA[
<p>I think this is true for projects beyond a certain complexity. I have 100% vibe coded projects with tens of thousands LOC, and haven't seen any real issues with fully automated maintenance. Will that approach work in every scenario, absolutely not, but the size and complexity of projects where it does is growing with each new model release.</p>
]]></description><pubDate>Sat, 13 Jun 2026 09:00:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48515088</link><dc:creator>MrScruff</dc:creator><comments>https://news.ycombinator.com/item?id=48515088</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48515088</guid></item></channel></rss>