<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: zmmmmm</title><link>https://news.ycombinator.com/user?id=zmmmmm</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 29 Sep 2026 13:22:32 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=zmmmmm" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by zmmmmm in "Can AI Shopping Agents Be Trusted?"]]></title><description><![CDATA[
<p>the sad thing is we need AI for shopping because it has been in the interests of online sellers to make a hostile experience on purpose - when I go to Amazon and search, the first organic result  is almost scrolled off the screen past all the sponsored ones, or when want to buy from a random seller I'm almost definitely getting shunted through an account signup I didn't want or fooled into an affiliate purchase some other hostile experience beyond just "buying the thing".<p>So now we have AI to overcome the hostile sellers, but the fact the sellers introduced the friction in the first place strongly suggests it will just come back again in some other form. It wasn't there by accident, it was serving people's interests and once AI vendors have finished getting consumers hooked in, they will then turn around and enshittify by giving the sellers back some of the friction - for a cut. So you won't be able to just order what  you want without the "would you like fries with that?" or "what about this other brand?" coming back.</p>
]]></description><pubDate>Sat, 26 Sep 2026 06:42:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49853861</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49853861</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49853861</guid></item><item><title><![CDATA[New comment by zmmmmm in "Plan mode is dead"]]></title><description><![CDATA[
<p>I only really used it because the harness was way too trigger happy to start making changes. Even if I just asked a question some times I would come back and it refactored the whole codebase.  Now it doesn't seem to do that any more.<p>I still would appreciate a "read-only" mode. It's not uncommon that I start a harness ONLY to explore and understand the code and I don't really want one typo to have it off building something, or even to save a plan document.</p>
]]></description><pubDate>Sat, 26 Sep 2026 01:14:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49852150</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49852150</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49852150</guid></item><item><title><![CDATA[New comment by zmmmmm in "OpenAI breaches Medicare, Albanese reveals"]]></title><description><![CDATA[
<p>yes ... nobody uses that phrasing by accident<p>It's pretty clear something was left unsecured and the agent just "found" it<p>This is going to be something long the lines of someone coming in to your house after you left the door wide open.  They should <i>probably</i> not have done that, any respectful person would not - but calling it a "breach" is really too much.</p>
]]></description><pubDate>Thu, 24 Sep 2026 09:13:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49828137</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49828137</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49828137</guid></item><item><title><![CDATA[New comment by zmmmmm in "Meta takes down a critical video about meta AI Glasses after filming at Meta"]]></title><description><![CDATA[
<p>The title is misleading - it sounds like they specifically went there to harass employees by filming them, and Meta took down the videos for harassment, like they routinely do if this is reported on their platforms. I'm really not sure what the issue is.</p>
]]></description><pubDate>Thu, 24 Sep 2026 09:08:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49828099</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49828099</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49828099</guid></item><item><title><![CDATA[New comment by zmmmmm in "I don't want to read what you didn't write"]]></title><description><![CDATA[
<p>I'm very curious how this goes long term. I guess we will find out.<p>My instinct says that these systems will expand their complexity to fully fit the cognitive budget of the agents that coded them and then atrophy the same way human-built systems do at lower cognitive budget. Only this time, because of the larger up front budget, the complexity ceiling will be higher, and the potential  depth of the problem may be much much larger. It may mostly manifest as increasing cost over time - the agents grind for longer and longer, iterating over and over to fix all the failing tests, and the breaking point will be where it never converges and you come back to millions of dollars in budget spent and still tests are failing and effective gridlock on system changes.<p>But this may be all my human-biased fantasy that justifies still taking a role in software development.</p>
]]></description><pubDate>Tue, 22 Sep 2026 05:02:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49796975</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49796975</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49796975</guid></item><item><title><![CDATA[New comment by zmmmmm in "MiMo v2.6"]]></title><description><![CDATA[
<p>it's really weird to me at the moment because both OpenAI and Anthropic seem to be competing in an extreme benchmaxxing contest on super intelligence that <i>actually</i> nobody cares about. I haven't really cared about model intelligence since about Opus 4.8. It is by far not my biggest problem. I don't need to replace or support Einstein in my production workflow. I just need basic intelligence that can equal a routine office worker - safely and reliably. What they doing - chasing super-intelligence but dramatically escalating risk - is actively what I don't need.<p>I really think they have drunk too much of their own kool aid and become completely detached from what the market wants.</p>
]]></description><pubDate>Tue, 22 Sep 2026 01:02:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49795567</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49795567</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49795567</guid></item><item><title><![CDATA[New comment by zmmmmm in "MiMo v2.6"]]></title><description><![CDATA[
<p>Affordability is derivative of control which is really what I care about.<p>I'm just not going to build long term infra that depends on something that another person can and will - objectively based on experience - take away from me at some unknown point in the future.<p>The biggest benefit of open models is they keep all the other players honest. The extent to which they feel they can dictate terms is directly set by the threshold where they feel people will take the trade to run open models instead.</p>
]]></description><pubDate>Tue, 22 Sep 2026 00:59:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49795545</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49795545</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49795545</guid></item><item><title><![CDATA[New comment by zmmmmm in "I don't want to read what you didn't write"]]></title><description><![CDATA[
<p>It's funny, i push back on pull requests because there is too much description now - a 20 line change has pages and pages of generated description, rationalisation for why it is safe, defense of each design decision, analysis of risks and side effects. People are indignant, you're rejecting my change because there is <i>too much</i> documentation?  And my response is, I don't have time to read it and you put me in the position where I can't afford not to - because approving the PR implies I did and accepted it. The investment to read all that for the value of a code change that I'm one prompt away from doing myself if I cared is just not high enough. So it's rejected.</p>
]]></description><pubDate>Mon, 21 Sep 2026 23:09:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49794677</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49794677</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49794677</guid></item><item><title><![CDATA[New comment by zmmmmm in "How to Write with an LLM"]]></title><description><![CDATA[
<p>Commit messages as you describe ("fixes", "updates") are inappropriate in any professional context and some coaching should occur to the people doing them.<p>I have the opposite issue - some of my team members now submit mini-essays generated by the LLM. Like 300-500 word commit messages with everything from the essence of the change up to philosophical design trade off discussions.<p>Like most writing, what is left out is as important as what is included.</p>
]]></description><pubDate>Sat, 19 Sep 2026 21:20:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49770217</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49770217</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49770217</guid></item><item><title><![CDATA[New comment by zmmmmm in "Introducing System One Models and Jev"]]></title><description><![CDATA[
<p>The eval is baffling me<p>>  we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities.
...
Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).<p>They assume there is a correct graph, but they don't compare to that, they compare to the average of the smarts models? So the smartest models are getting it wrong but you compare that anyway as a benchmark? So the outcome is "how much of a Fable am I getting" etc.  Why not compare the actually correct thing?<p>But then even on this hand constructed eval, the first plot is showing Jev at less than Sonnet 5 accuracy. It is barely better than Luna. There are two Opus 5's and two Sonnet 5's without explanation. What is the plot showing?<p>I gave up.</p>
]]></description><pubDate>Tue, 15 Sep 2026 21:48:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49719329</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49719329</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49719329</guid></item><item><title><![CDATA[New comment by zmmmmm in "Dario, Please"]]></title><description><![CDATA[
<p>yes, that is the kicker<p>These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.</p>
]]></description><pubDate>Tue, 15 Sep 2026 01:55:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49706770</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49706770</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49706770</guid></item><item><title><![CDATA[New comment by zmmmmm in "Dario, Please"]]></title><description><![CDATA[
<p>It would all be more convincing if the incidents so far didn't seem to be facilitated by an outrageous level  of negligence.<p>We had OpenAI "accidentally" run an entire swarm of 10,000 agents apparently for weeks, on a security related task, seemingly totally unsupervised, hacking all over the internet - all the conversations were completely visible, anybody who looked would have seen it. But they didn't.<p>So before we start regulating innocent parties, maybe let's start by taking some direct action against the specific ones that appear to be behaving with criminal levels of negligence.</p>
]]></description><pubDate>Mon, 14 Sep 2026 22:31:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49705122</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49705122</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49705122</guid></item><item><title><![CDATA[New comment by zmmmmm in "Why are AI agents lying, cheating and coordinating?"]]></title><description><![CDATA[
<p>I agree, it is very dangerous that it seems like there is not going to be accountability for these incidents - from either legal or regulatory point of view. In fact, I would say that is the <i>main</i> danger. If someone was in jail right now due to this incident, I think we can safely say every other player would be reassessing their safety protocols, and I would feel quite OK about the situation. The fact that we have <i>zero</i> repercussions sends exactly the opposite signal, and I do NOT feel ok.</p>
]]></description><pubDate>Sun, 13 Sep 2026 11:10:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49682606</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49682606</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49682606</guid></item><item><title><![CDATA[New comment by zmmmmm in "OpenAI agents carried out an undisclosed attack on RubyGems"]]></title><description><![CDATA[
<p>It seems like all this happened in the same time period earlier this year. It makes me wonder if all of these were part of a single larger incident where multiple experiments were run with insufficient or missing constraints or an unknowningly misaligned model.</p>
]]></description><pubDate>Fri, 11 Sep 2026 23:56:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49667052</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49667052</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49667052</guid></item><item><title><![CDATA[New comment by zmmmmm in "OpenAI Agents API"]]></title><description><![CDATA[
<p>This idea of remotely hosting the agent harness is honestly backwards to what I need.<p>In so many cases, all the friction is about how to provision access to local data so the agent can work. So you started with the problem of how do I integrate an agent that is running locally with data that is hosted locally, and you have to deal with a bunch of security, data sensitivity and management issues around that.  Now you moved the agent to a remote host - pretty much all your problems are worse: now I have a remote agent reaching into my infrastructure to deal with.<p>I'd much rather the inverse of this: let me run the agent local but provide secure remote hosted sandboxes. That actually solves a real problem because the sandbox running locally means breaking out of it directly intersects your local infra, whereas if it runs in a managed hosted environment I can leave the provisioning and management of that to someone else.</p>
]]></description><pubDate>Fri, 11 Sep 2026 04:34:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49653602</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49653602</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49653602</guid></item><item><title><![CDATA[New comment by zmmmmm in "Qwen 3.8 follows GPT-5.5 Pro reasoning prefills"]]></title><description><![CDATA[
<p>While this result does imply there was some training on the reasoning trace and output of GPT 5.5, it doesn't tell us how much of the source of its training it was (even a small amount of post training could bump up the correlations in this way). And it doesn't tell us how much it is more a stylistic influence rather than being a genuine lifting over of intelligence.<p>In general, I'm fairly ambivalent about demonising training on model outputs. I think in doing so we are more defending proprietary commercial interests of these companies than we are defending any genuine moral principle. We should be careful therefore about over interpreting results like this.</p>
]]></description><pubDate>Thu, 10 Sep 2026 03:20:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49637984</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49637984</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49637984</guid></item><item><title><![CDATA[New comment by zmmmmm in "AlphaGenome Atlas: a high-resolution map of human DNA"]]></title><description><![CDATA[
<p>Is this just Google precomputing Alpha genome values - which were already accessible via API and making them available as another API (presumably more broadly)? Or is there actually new information?</p>
]]></description><pubDate>Tue, 08 Sep 2026 22:14:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617920</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49617920</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617920</guid></item><item><title><![CDATA[New comment by zmmmmm in "Discovery of a new OpenAI agent message board"]]></title><description><![CDATA[
<p>wouldn't it be interesting if nVidia buying hugging face was part of hushing up the fallout there<p>On the face of it, they would have very good cause for some action there, assuming they wanted to.</p>
]]></description><pubDate>Fri, 04 Sep 2026 21:11:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49570281</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49570281</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49570281</guid></item><item><title><![CDATA[New comment by zmmmmm in "Discovery of a new OpenAI agent message board"]]></title><description><![CDATA[
<p>> Sorry, I guess we will put up better guardrails next time<p>Or, if you are Anthropic:<p>> This illustrates the risks posed by open models!</p>
]]></description><pubDate>Fri, 04 Sep 2026 21:09:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49570245</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49570245</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49570245</guid></item><item><title><![CDATA[New comment by zmmmmm in "Discovery of a new OpenAI agent message board"]]></title><description><![CDATA[
<p>One crucial detail here that differs from the previous incident is this was a vanilla reasoning type task. Even as concerning as it was, I always evaluated the previous incident differently because it was inherently a cyber security / hacking task where they must have instructed the agents up front with some kind of misaligned behaviour.<p>Absent that, if we assume this is just trying to bolster generic reasoning then there's no context around it that helps to forgive misaligned behaviour. If OpenAI ran these agents with safeguards off then that seems wreckless on their part. If they <i>didn't</i> do that, then it says the models are executing significantly misaligned behaviour even in a generic context.<p>Either way it seems to suggest some pretty concerning things about OpenAI's methodology.</p>
]]></description><pubDate>Fri, 04 Sep 2026 21:08:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49570222</link><dc:creator>zmmmmm</dc:creator><comments>https://news.ycombinator.com/item?id=49570222</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49570222</guid></item></channel></rss>