<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: ekidd</title><link>https://news.ycombinator.com/user?id=ekidd</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 28 Jul 2026 08:12:45 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=ekidd" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by ekidd in "What is happening to jobs? Separating AI hype from reality"]]></title><description><![CDATA[
<p>> My experience so far, has been that it's barely a junior dev.<p>At least on greenfield projects, Fable is more like having an endless succession of fly-by senior devs who lock themselves in an office for a week, and who come back with very reasonable code that I then need to maintain somehow.<p>For longer-term maintenance, I am actually slowly warming to Sonnet 4.5-era models (so October 2025, right before the Opus revolution). They need to be given clear instructions and watched carefully. But since they force a human to stay in the loop, you don't have the institutional knowledge loss I see in some Opus projects, or the code that was one-shot with no human interaction at all that's a constant temptation in Fable projects.<p>And yeah, the bit rot is <i>painfully</i> real once the humans step back too far. I've been dealing with a compelling prototype that someone built, and that stakeholders love (for good reasons). But it had to be put on a tech debt repayment plan for a couple of months.</p>
]]></description><pubDate>Sun, 26 Jul 2026 14:36:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49058640</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49058640</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49058640</guid></item><item><title><![CDATA[New comment by ekidd in "The new rules of context engineering for Claude 5 generation models"]]></title><description><![CDATA[
<p>When I learned French as an adult, I pushed myself heavily into "forced immersion", avoiding English completely for long periods of time as my French slowly built itself up. This almost entirely silenced my inner monologue in English for those periods of time, and left me with a toddler's ability to think in French. In this comparative silence, it was much easier to observe my non-verbal thought processes moving around "behind" the scarce words.<p>Then my French inner monologue got good enough that I could mostly think in French, especially when I was in a French-speaking environment. One fascinating detail was that after switching from a French-speaking environment to an English-speaking one, I would actually spontaneously <i>translate from French to English</i> for about 15 minutes until my brain switched back.<p>So it seems obvious to me that it's possible to suppress or at least severely impoverish the language of thought, that other "layers" of thought exist besides the words, and that it's even possible to change the actual language of verbal thought.<p>Also, something which at least <i>some</i> other people in the HN crowd might recognize: When I'm deepest in the zone programming and refactoring, I tend to work with a lot of half articulated concepts I can't put into words. You know how people talk about "code smells"? That isn't a literal smell for me, but it's generally a non-verbal sense that a pattern is wrong.</p>
]]></description><pubDate>Sun, 26 Jul 2026 12:04:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49057280</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49057280</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49057280</guid></item><item><title><![CDATA[New comment by ekidd in "Nvidia, Microsoft, Meta warn against overregulating open-weight models"]]></title><description><![CDATA[
<p>DeepSeek has published some really good papers. Lately they're pushing really hard for dramatically cheaper serving costs.</p>
]]></description><pubDate>Fri, 24 Jul 2026 21:57:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49042102</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49042102</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49042102</guid></item><item><title><![CDATA[New comment by ekidd in "The arguments against open source AI are bad"]]></title><description><![CDATA[
<p>> <i>Especially for point #1 I don't think we've established that - we've been given information by a private company that makes their tooling look extremely valuable which may be true and genuine or may just be yet another doomday statement to bolster their valuation.</i><p>Many of the details of the attack on Huggingface were reported by <i>them</i> before they knew who was attacking. So no, OpenAI is not the only source here. It was a pretty impressive example of an APT-style attack just from their end.<p>"Our model is powerful enough to commit multiple felonies (and we can't stop it)" is "marketing," I suppose.</p>
]]></description><pubDate>Thu, 23 Jul 2026 21:40:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49028414</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49028414</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49028414</guid></item><item><title><![CDATA[New comment by ekidd in "The arguments against open source AI are bad"]]></title><description><![CDATA[
<p>Everyone <i>always</i> says things like "It's just marketing".<p>But seriously, I highly recommend reading the published details of the Huggingface breach. The model found and chained multiple zero days. To escape, it punched a hole in a commercial package repository proxy (sort of like an npm mirror) using a previously unknown bug. From there, it needed to move laterally through OpenAI internal systems to actually reach a network. To attack Huggingface, it used multiple new zero-day security holes plus credentials that it stole. The model also had sufficient long-term planning and agent-management capabilities to maintain focus on a sustained attack.<p>Any attack which requires weaponizing multiple zero-days and maintaining state for an ongoing attack like this is (1) beyond the "attention span" of publicly available models, and (2) pretty much the definition of an Advanced Persistent Threat.<p>I assume that these Galaxy-class models are not available to public because:<p>1. They're almost certainly too expensive to serve at scale. These are the models OpenAI uses to solve famous math problems for headlines, not actual viable products yet.<p>2. OpenAI doesn't know how to keep them from going off the rails like this. Remember, the Huggingface attack happened because the model was asked to do a cybersecurity benchmark. It escaped containment and broke into Huggingface to <i>steal an answer key.</i> Very few corporations want the liability associated with models that act like this.<p>> <i>Security through obscurity isn't tenable anymore.</i><p>I absolutely agree with this. The "only way out is through" with computer security, and I expect it to be an ugly few years.</p>
]]></description><pubDate>Thu, 23 Jul 2026 19:40:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49026974</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49026974</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49026974</guid></item><item><title><![CDATA[New comment by ekidd in "The arguments against open source AI are bad"]]></title><description><![CDATA[
<p>> <i>If AI were to become a super weapon why should I trust a private company to own it?</i><p>We have just recently established that:<p>1. OpenAI's internal "Galaxy" model is fully capable of functioning as what security people refer to as an "Advanced Persistent Threat." The published details of the recent sandbox escape and Hugging Face attack involved chaining multiple unknown zero-days at various stages of the attack, and executing an ongoing adaptive attack. This is previously a state-level ability, or at least something you'd expect from people on the CTF leaderboards.<p>2. OpenAI is clearly incapable of <i>controlling</i> their in-house models. This is the second time Galaxy-class models are known to have breached containment and done bad stuff.<p>It is highly likely that versions of these offensive abilities will be widely available within a year or so. At which point I expect widespread incidents similar to what happened to Hugging Face. We aren't ready for this.<p>But yes, if AI becomes an even more dangerous weapon that <i>that</i>, it's time to start asking questions like "What the hell are we doing, anyway?"</p>
]]></description><pubDate>Thu, 23 Jul 2026 19:14:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49026655</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49026655</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49026655</guid></item><item><title><![CDATA[New comment by ekidd in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>Oh, to be clear, I don't think that <i>anything</i> about this overall situation is even slightly OK.</p>
]]></description><pubDate>Wed, 22 Jul 2026 05:14:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49002136</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49002136</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49002136</guid></item><item><title><![CDATA[New comment by ekidd in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>Apparently, the attacker in the Hugging Face case was reported to be an internal OpenAI model trying to break into HF and steal the answers to cybersecurity benchmarks: <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="nofollow">https://openai.com/index/hugging-face-model-evaluation-secur...</a> It really doesn't matter what restrictions are placed on public use of models if the attacking models are internal models at the AI labs themselves.<p>So if the AI labs are literally running rogue models breaking into other organizations' servers, then yes, I am OK with those organizations self-hosting Chinese models for defensive use.</p>
]]></description><pubDate>Wed, 22 Jul 2026 00:22:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49000212</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=49000212</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49000212</guid></item><item><title><![CDATA[New comment by ekidd in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>> <i>For small models (which are probably distilled from their big ones) you can serve them economically all the time and not hemorrhage money.</i><p>For smaller models, you're competing with DeepSeek V4 Flash. (Which I think is a 284B A13B?) Subjectively, this feels about as smart as Sonnet 4.5, give or take. And it costs $0.09/$0.18 on Open Router, compared to $1/$5 for the latest Claude Haiku. See <a href="https://openrouter.ai/deepseek/deepseek-v4-flash#providers" rel="nofollow">https://openrouter.ai/deepseek/deepseek-v4-flash#providers</a> The developer antirez of Redis fame uses this as a local coding model.<p>DeepSeek did some extremely clever research on hybrid attention to get the prices that low, reducing per-user context cache sizes dramatically.<p>So, no, when it comes to low-price models, the US models probably can't sustain their current margins there, either.</p>
]]></description><pubDate>Tue, 21 Jul 2026 22:06:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999006</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48999006</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999006</guid></item><item><title><![CDATA[New comment by ekidd in "Annoying and alarming things about OpenCode"]]></title><description><![CDATA[
<p>Those are Claude Code security issues? And Claude Code is a gigantic vibe-coded dumpster fire with more bugs than an ant hill?<p>pi-agent famously doesn't even <i>try</i> to provide a sandbox, so obviously it can't have sandbox security bugs!<p>I actually think that this is the right model: The agent should not have access to anything it doesn't actually need: User files outside the work tree (except for maybe some allow-listed dotfiles), network access, real credentials, etc. Lock it down tight, and reduce exposure on multiple parts of the "Lethal Trifecta".</p>
]]></description><pubDate>Mon, 20 Jul 2026 14:56:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48979745</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48979745</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48979745</guid></item><item><title><![CDATA[New comment by ekidd in "Annoying and alarming things about OpenCode"]]></title><description><![CDATA[
<p>It is a pretty catastrophically <i>dumb</i> CVE, of the sort that makes me not want to allow OpenCode anywhere near any of my machines in the future. It was basically "RCE as a service", not some subtle bug.<p>Personally, I run pi-agent in a custom sandbox based on bwrap, with an internet proxy. This mostly limits the blast radius to one source tree and one git checkout. And I don't give it push/pull permission. Local models are fun in a "closely supervised but slightly clueless minion" sort of way.</p>
]]></description><pubDate>Mon, 20 Jul 2026 14:12:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48979150</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48979150</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48979150</guid></item><item><title><![CDATA[New comment by ekidd in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>AI subscription pricing was fine when it was $100/month for some opaque 5 hour token budget I don't think I ever used, not even that one day where I coded for 14 hours non-stop using Fable. But like most people with low token usage, I had a human in the loop and and I didn't use workflows with swarms of agents.<p>Now, of course, the plan is to remove Fable from the subscription. To paraphrase Darth Vader, they have altered the deal. Pray they do not alter it further.</p>
]]></description><pubDate>Sat, 18 Jul 2026 21:37:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=48962680</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48962680</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48962680</guid></item><item><title><![CDATA[New comment by ekidd in "Why write code in 2026"]]></title><description><![CDATA[
<p>I have written assembly for about 5 different processors, including 65C02s, 680x0s, cute little DSP3210s that managed the CPU cache manually, utterly cursed TI320C40s, and (of course) a bit of Intel. WebAssembly is simpler than some of these architectures in <i>some</i> ways.<p>But it's not <i>that</i> much simpler. And once you add the WASM GC stuff, WebAssembly gets weird. It's a Harvard architecture with separate value memory, linear memory, GC memory <i>and</i> "tables", all accessible in completely different ways, with a weird mandatory type system (especially for the GC stuff). And the docs are often terrible. And yes, I've also written WebAssembly by hand.<p>So yes, I would, overall, classify WebAssembly as "slightly easier". But not dramatically so. And the training data for actually writing non-trivial things by hand isn't <i>that</i> great, not compared to something like Intel assembly.<p>(Don't talk to me about TI320C40 assembly. If Fable can one-shot a Prolog interpreter written using <i>that</i> without finding a reference manual, it's time to hang up my hat and learn to make goat cheese.)</p>
]]></description><pubDate>Mon, 13 Jul 2026 14:50:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48893612</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48893612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48893612</guid></item><item><title><![CDATA[New comment by ekidd in "Why write code in 2026"]]></title><description><![CDATA[
<p>Fable wrote a pretty decent test suite covering typical Prolog programs, including things with non-trivial execution patterns like "append" that do complex backtracking on multiple branches. And I've run a modest number of test queries by hand, trying a few things. It's entirely possible that there's a bug there somewhere. But it's better than <i>I</i> would have done on the first try, implementing a Prolog interpreter in assembly. And I've worked on actual production compilers a few times.<p>I am honestly <i>not happy</i> about the way that models can now just take what should have be a fun multi-weekend project and knock out in a couple of hours. But I'm not going to pretend that Fable is stupid, or that it did a bad job on any of the test projects I gave it. It struggles more on big, messy real-world code bases, absolutely.</p>
]]></description><pubDate>Mon, 13 Jul 2026 14:36:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=48893397</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48893397</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48893397</guid></item><item><title><![CDATA[New comment by ekidd in "Why write code in 2026"]]></title><description><![CDATA[
<p>Claude is perfectly capable of writing assembly. Here's a working (basic) Prolog interpreter that Claude Fable 5 wrote in WebAssembly in 61 minutes for $16.75 in token costs: <a href="https://github.com/emk/fable-wasm-prolog/blob/main/prolog.wat" rel="nofollow">https://github.com/emk/fable-wasm-prolog/blob/main/prolog.wa...</a><p>WebAssembly is slightly easier than real assembly, but here Fable used WASM GC extensions, which are poorly documented and not yet super common.<p>Fable didn't even need to debug it; I believe essentially all the assembly worked correctly on the first try.<p>I have <i>feelings</i> about this, but I'm not pretending it isn't real.</p>
]]></description><pubDate>Sun, 12 Jul 2026 21:10:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48884828</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48884828</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48884828</guid></item><item><title><![CDATA[New comment by ekidd in "I love LLMs, I hate hype"]]></title><description><![CDATA[
<p>You can maybe run a local Sonnet-4.5-ish-level model (<i>sort</i> of) for less than the price of a new car, even at current massively inflated prices for fast RAM. This is probably not what you were looking for. But it's there. You could share one server between multiple developers. Maybe make a little AI co-op or something, with a pair of RTX Pro 6000 cards?<p>Also, DeepSeek V4 Pro is cheap via any commodity API, and DeepSeek V4 Flash is essentially free at API prices like $0.09/M, $0.18/M out. This is generally not subsidized.<p>For a more practical local setup, Qwen3.6 27B on a used Nvidia 3090 (US$1300) or two is surprisingly nice. It needs clear instructions and you can't use it for hands-off vibecoding, but it's actually quite reasonable for hands-on programmers.</p>
]]></description><pubDate>Sun, 12 Jul 2026 20:22:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48884391</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48884391</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48884391</guid></item><item><title><![CDATA[New comment by ekidd in "Datacentres drive up big tech's carbon emissions to a third of those of France"]]></title><description><![CDATA[
<p>GLM 5.2 isn't <i>quite</i> modern Opus tier, as seen in this comparison where Opus 4.5 scores 4/5 on some coding tasks where GLM 5.2 scores 0/5: <a href="https://www.tryai.dev/blog/gpt-5.6-build-off-12-models" rel="nofollow">https://www.tryai.dev/blog/gpt-5.6-build-off-12-models</a> But yes, GLM 5.2 is cheap.<p>But the real standout on price is DeepSeek V4 Flash, which competes, more or less, with models in between Sonnet and Haiku. From third-party providers, it costs around $0.09/M, $0.18/M out, compared to $3M/in, $15M/out for Sonnet and $1M/in, $5M/out for Haiku. To get the price of DSv4 (Flash and Pro) so low, DeepSeek did a lot of innovative optimization work that will likely show up in other open weight models in the future.</p>
]]></description><pubDate>Sun, 12 Jul 2026 13:09:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48880915</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48880915</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48880915</guid></item><item><title><![CDATA[New comment by ekidd in "Amber the programming language compiled to Bash/Ksh/Zsh"]]></title><description><![CDATA[
<p>A really common use case is install/bootstrap/setup scripts. You know, those sketchy curl-to-bash things, or cloud-init scripts, or whatever you run to set up your actual higher level tools like Python.</p>
]]></description><pubDate>Sun, 12 Jul 2026 00:02:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48877007</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48877007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48877007</guid></item><item><title><![CDATA[New comment by ekidd in "Who manages the agents?"]]></title><description><![CDATA[
<p>> <i>I don't take 3 as a given. There's just too much going on in the space for one cloistered company to control it all and be in control of it.</i><p>Yeah, in my comment, I was assuming the publicly-stated goals of the labs actually came true. I assumed that they achieved true AGI (defined as "about as smart as Fable, except it can manage long-term tasks as well as a smart human, too"), and that this would cause widespread unemployment and fully-automated production. And furthermore, I assumed that once they could use AGI to automate <i>more</i> AI research, they could use that to make even smarter models. The labs call this "recursive self-improvement" (RSI) and like to insist it will happen Real Soon Now.<p>Given <i>this</i> chain of events come true, then I think there's an excellent chance that the resulting models will never be sold via an API. Given those assumptions, the labs would probably make more money by keeping all the compute to themselves and just ordering the AI to start and manage businesses. Imagine Anthropic having an in-house OpenClaw that could plan and run a successful startup with no human input besides an annual "board meeting" with the humans.<p>This isn't the only possible outcome, of course.<p>- Maybe the labs are wrong and progress stalls out well short of superintelligence. This would be nice!<p>- Maybe true AGI is only requires one or two clever algorithmic tricks beyond what we have now. In this case, training costs might go <i>down</i>, and superhuman models might become widespread. I suspect that possible future would be extra weird.</p>
]]></description><pubDate>Sat, 11 Jul 2026 20:54:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48875775</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48875775</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48875775</guid></item><item><title><![CDATA[New comment by ekidd in "Who manages the agents?"]]></title><description><![CDATA[
<p>The author is completely right about the <i>AI Lab's</i> promised vision of the world: They claim to want to create superhuman intelligence, which will produce vast abundance. But superhuman intelligence would be extremely dangerous, so it needs to be controlled by a tiny "priesthood" of trusted people, or somehow designed so that the superhuman intelligence could be trusted. (We have no idea how to do that.)<p>But the author's vision is <i>also</i> suspect, if you assume that the models will become much more intelligent:<p>1. Hypothetically, we can't give every human their own personal SkyNet to command. That would, uh, probably end very badly. If everyone gets an agent, those agents can't be <i>too</i> capable?<p>2. If you do somehow build a model that's much smarter than you, what do you contribute by managing it? How many people here have ever worked for a well-intentioned manager who couldn't understand the people they managed? So in this scenario, human management would be mostly displaced by agent management. Most companies could lay almost everyone off and let the agents manage each other. We only need humans to manage models now because the models are still pretty broken.<p>3. If we create models that can genuinely replace humans at almost any task, you won't be able to buy those on the API. At that point, the billionaires and the politicians wouldn't <i>need</i> human workers any more, because everything can be done better using their pet agents. Just have the robots build stuff for the billionaires directly. And if any of the former human peons get upset about being locked out of the economy to starve, then have the agents pilot the drones, too.<p>Basically, almost none of the people imagining a future of superhuman intelligences have actually though through how it would actually work in the real world. We're going to spend trillions of dollars and vast amounts of resources chasing the goal of making ordinary humans obsolete. Now, that goal might be unobtainable, I hope. But I'm deeply alarmed at how much we're spending pursuing it.</p>
]]></description><pubDate>Sat, 11 Jul 2026 18:46:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=48874574</link><dc:creator>ekidd</dc:creator><comments>https://news.ycombinator.com/item?id=48874574</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48874574</guid></item></channel></rss>