<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: 827a</title><link>https://news.ycombinator.com/user?id=827a</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 30 Aug 2026 00:34:51 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=827a" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by 827a in "Managing AI Coding Costs at Scale"]]></title><description><![CDATA[
<p>I'm probably between $50-$200/day depending on the day; we also have effectively unlimited budget, though a lot of that is because Azure gives startups $150,000 in credits for 2 years, which we've wired up to a LiteLLM gateway & OpenCode. Without that I think our appetite would be more around $400/month/employee.<p>A lot of my high costs is because I just throw Sol at everything. If I were more selective and brought in Luna or v4 Flash every once in a while, I think I'd be more like ~$400/month. That's why I'm not aligned with the notion that "tokens are subsidized so that's why people are using so much": its not that I'll have to adjust to using less, its just that I'd need to think before I prompt a bit and be more judicious. I could easily see my raw token counts doubling or tripling in the coming months. I don't think that will change as subsidization subsides; though maybe lab revenue will; intelligence per dollar is getting cheaper every week. Its solely a function of adaptation to process, which takes time.<p>The productivity gains per token are the single most asymmetrical thing I've ever seen in engineering. The engineers on our team are pretty effective with tokens; easily that 2x-4x output as you're seeing, spending $20-$200/day. Some of our security folks have also started contributing more-and-more code, and they're on the other side: they'll spend hundreds a day running in circles, eventually producing these +/-30k loc pull requests that take ages to get merged and are littered with issues. They weren't writing much code before, so arguably they're more productive by some multiplier greater than 1, but I think the drag on the rest of the team, and potential issues with what they produce, has overall created a net-negative situation. Inversely, some other company functions have produced a few one-off websites for things like sales processes, and <i>those</i> have been a huge win. The asymmetry is wild. There's almost a valley of incoming skill where if you know nothing about code, you'll leverage it well; if you know just a little bit, it makes you super dangerous; if you know a lot, you're the biggest winner. Really difficult situation to navigate.</p>
]]></description><pubDate>Sat, 08 Aug 2026 03:32:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49218677</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49218677</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49218677</guid></item><item><title><![CDATA[New comment by 827a in "When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation"]]></title><description><![CDATA[
<p>Your issue, I believe, is that you seem to believe capabilities are measured along one axis. This is natural to believe because it is representative of how the models have evolved up to this point, and thus it is also what many AGI-pilled people believe.<p>Critically, you did not quote the most important part of my sentence: "useful progress will probably slow down and become more linear starting in Q4"; your omission of those words is why I believe you don't understand what I'm saying; you didn't find it important to make your point, so you omitted it, when actually it is critical to the entire assertion. You can read my third paragraph, if you wish, to understand why it is important, instead of just stopping at the first word you disagree with and hitting the "Submit Comment" button.</p>
]]></description><pubDate>Wed, 05 Aug 2026 16:14:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49184870</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49184870</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49184870</guid></item><item><title><![CDATA[New comment by 827a in "When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation"]]></title><description><![CDATA[
<p>Coding is an extremely verifiable and loopable task, like math (in fact, all of the math that these models has done has been through the lens of Lean, which is itself just coding). I am talking about their capabilities in tasks that are more general, the execution of which represent the vast majority of economic value generation in the world.</p>
]]></description><pubDate>Wed, 05 Aug 2026 13:49:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49182884</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49182884</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49182884</guid></item><item><title><![CDATA[New comment by 827a in "When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation"]]></title><description><![CDATA[
<p>I'm very bullish on AI, but I still feel that we pretty much plateaued at Opus 4.6 and everything since then has been in the domain of "extremely verifiable and loopable tasks" (math), benchmarkmaxxing, and harness improvements. Which are fine to good things, but I think its very reasonable at this point to start asking questions about when we'll see progress in more general domains. The readability of AI output, for example, has nosedived as they've gotten more intelligent, which makes the frontier models difficult to use even for things like writing emails.<p>In that sense, the frontier models are going to quickly blaze past any semblance of usefulness to humans, while every once in a while we get a news drop like "GPT-7 solved some crazy math problem" or "it invented some new awesome drug"; meanwhile what most people will use will be smaller, more human-specialized models, maybe distilled from those frontier models, that take much longer to iterate on because they rely on large amounts of human feedback in the domain they're specialized for. In other words, useful progress will probably slow down and become more linear starting in Q4, bounded by the rate at which the humans paying for it say "yes this is a good react website".<p>(By the way: I earnestly do categorize "inventing a new drug" as non-useful AI progress, counter-intuitively. The drug industry has more ideas for drugs than they know what to do with; "useful progress" is, after the idea is made, validating that it works in humans and doesn't kill the human, and productionizing it. AI will help with this and does, but I have substantial doubt that we'll ever see the drug pipeline speed up to, like, a year from idea to prescription. <i>That</i> would be useful progress, which unfortunately many AI pilled hypermaxers conveniently forget. The invention of a promising new drug, or the solution to an arcane set theory problem, are cherries that, through the diligent labor of humans and AI, may become useful, but progress is rarely made by the lone intellect having an a-ha moment.)</p>
]]></description><pubDate>Wed, 05 Aug 2026 03:19:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49178215</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49178215</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49178215</guid></item><item><title><![CDATA[New comment by 827a in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>Totally untrue. Luna and Sonnet 5 are very comparable: <a href="https://artificialanalysis.ai/#intelligence" rel="nofollow">https://artificialanalysis.ai/#intelligence</a><p>Luna is an extremely strong model.</p>
]]></description><pubDate>Thu, 30 Jul 2026 18:07:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49113532</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49113532</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49113532</guid></item><item><title><![CDATA[New comment by 827a in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>Vera Rubin will be hitting racks very soon, and this is purported to have a 10x improvement in token throughput per megawatt. Of course, old chips don't get replaced with new chips overnight, but I don't think we're anywhere near the floor yet.</p>
]]></description><pubDate>Thu, 30 Jul 2026 18:06:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49113514</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49113514</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49113514</guid></item><item><title><![CDATA[New comment by 827a in "Passkeys were invented by engineers with zero understanding of consumer brain"]]></title><description><![CDATA[
<p>Passkeys are so, so, so bad. One of the worst things our industry invented. The sooner sites start leaving them on the wayside and just go back to TOTP, SMS, and Email codes/links, the better. These work. We solved auth. Its fine.</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:21:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010183</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=49010183</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010183</guid></item><item><title><![CDATA[New comment by 827a in "Cursor 0day: When Full Disclosure Becomes the Only Protection Left"]]></title><description><![CDATA[
<p>They should definitely fix it, but that's mostly because its an "unnecessary autoplay" so to speak. There's plenty of "necessary autoplays" out there, and AI is going to add more and more every day, because that's where productivity comes from. But, why Cursor would ever need to execute the git binary in your project directory is beyond me; very clearly a bug.<p>Their ignorance of the bug report is also very clear and concerning negligence.<p>But I think simultaneously, the security team is making a mountain out of a molehill. This is a classic thing security teams love doing; everything is military defcon P0. So, its important to check them regularly, and remind them that the most secure system is no system; they are but one part of a greater ecosystem.</p>
]]></description><pubDate>Wed, 15 Jul 2026 04:05:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48916178</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48916178</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48916178</guid></item><item><title><![CDATA[New comment by 827a in "Cursor 0day: When Full Disclosure Becomes the Only Protection Left"]]></title><description><![CDATA[
<p>Frankly, if you git clone a compromised repository, I'm not sure that a vulnerability of the class "compromised code in that repository will be executed" is all that major a concern. There are plenty of IDEs that will go autonomously run npm installs (with post-install scripts) for you when they detect a package.json. This isn't all that different than that.<p>They could throw up a warning like "do you trust this repository" oh wait they already do, and no one cares. Security is hard. Ultimately if you have compromised code on your machine, all bets are off.</p>
]]></description><pubDate>Tue, 14 Jul 2026 21:21:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=48913105</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48913105</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48913105</guid></item><item><title><![CDATA[New comment by 827a in "The Supreme Court Just Lit a Fuse Under Flock's License Plate Camera Empire"]]></title><description><![CDATA[
<p>The constitution does not really care about scale, though, and that’s my point. It’s a reason why the legislature should care about Flock, but not why the judicial should.</p>
]]></description><pubDate>Sat, 11 Jul 2026 13:12:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48871786</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48871786</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48871786</guid></item><item><title><![CDATA[New comment by 827a in "The Supreme Court Just Lit a Fuse Under Flock's License Plate Camera Empire"]]></title><description><![CDATA[
<p>The constitution does not really care about scale, though, and that’s my point. It’s a reason why the legislature should care about Flock, but not why the judicial should.</p>
]]></description><pubDate>Sat, 11 Jul 2026 13:11:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=48871779</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48871779</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48871779</guid></item><item><title><![CDATA[New comment by 827a in "The Supreme Court Just Lit a Fuse Under Flock's License Plate Camera Empire"]]></title><description><![CDATA[
<p>The argument I struggle to get around and would love to hear a counter-argument to: Let's say a local police department hired 175 police officers, each being told "Go stand on this particular intersection with a pad of paper and write down every license plate you see". This would be a stupid use of resources, but is not outside the realm of something a well-funded police department could do. Every night they take their reports back to HQ, and file them away.<p>This is a modestly different situation than one concerning warrantless tracking of phone locations, if for no other reason than my phone oftentimes in my pocket. It is not always visible to onlooking bystanders. And even if it isn't, externally there is no reliably way to differentiate one iPhone from another. In comparison: license plates, when in public, are always visible, and very easy to discern from one-another (different state-unique numbers); so in my mind the expectation of privacy is far lower.<p>I abhor what Flock does, but I'm not sure I see a constitutional argument for why what they do is unconstitutional.</p>
]]></description><pubDate>Mon, 06 Jul 2026 19:06:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48809079</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48809079</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48809079</guid></item><item><title><![CDATA[New comment by 827a in "AMD Ryzen AI Halo – $4k AI Dev Kit"]]></title><description><![CDATA[
<p>Microcenter has them marked at $4500 right now (that's with the 4TB SSD) [1]. I suspect it comes down to what you're using it for; if you're looking for a general purpose computer that's also solid at AI, the AMD machine is better. But if you want the best possible AI machine at below $5k... actually you should probably just buy an RTX 3090 or 5090. But if the 128gb of memory is critical, then yeah DGX Spark is it.<p>[1] <a href="https://www.microcenter.com/product/699008/nvidia-dgx-spark" rel="nofollow">https://www.microcenter.com/product/699008/nvidia-dgx-spark</a></p>
]]></description><pubDate>Mon, 06 Jul 2026 18:49:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=48808849</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48808849</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48808849</guid></item><item><title><![CDATA[New comment by 827a in "AMD Ryzen AI Halo – $4k AI Dev Kit"]]></title><description><![CDATA[
<p>Framework, weirdly, overcharges considerably for their SSDs. You can currently get a Samsung 990 Pro 2TB on Amazon for $390; Framework charges $625 for the Sandisk 850x 2TB, which has similar performance (and is being sold on Amazon for $530).<p>If you DIY your own SSD, you can spec a Framework Desktop for below $4k; but not much below. Roughly the same price.</p>
]]></description><pubDate>Mon, 06 Jul 2026 18:43:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=48808764</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48808764</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48808764</guid></item><item><title><![CDATA[New comment by 827a in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>Tbh we'll see what using it looks like, but the reasoning/cost charts do not look promising. It seems like the only useful reasoning level for Sonnet 5 is Low; medium might trade blows at price/performance with Opus, but anything beyond that Opus is Just Better.<p>I struggle to understand where this model fits in. If I need a cheap model for simple stuff (like, summarizing an email); I'd go Haiku (actually, I'd go Deepseek v4 Flash, but you catch my drift). I just can't think of many tasks where I'm like "yeah let me reach for Sonnet Low Reasoning so I can save a dollar but also seriously run the risk of it failing"; I'd just reach for Opus Low.</p>
]]></description><pubDate>Tue, 30 Jun 2026 20:14:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48738578</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48738578</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48738578</guid></item><item><title><![CDATA[New comment by 827a in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>Why are you comparing xhigh reasoning between Sonnet and Opus? Of course Sonnet xhigh is cheaper than Opus xhigh, but that isn't the point; the point is that at e.g. 80% accuracy on Opus costs ~$0.45 (medium reasoning) whereas on Sonnet it costs ~$0.52 (xhigh/max reasoning).</p>
]]></description><pubDate>Tue, 30 Jun 2026 20:10:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48738536</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48738536</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48738536</guid></item><item><title><![CDATA[New comment by 827a in "Claude Code is steganographically marking requests"]]></title><description><![CDATA[
<p>This seems really, really stupid. Similar to the weird Zig runtime signature thing from a few months ago ago, it was bound to be discovered, quickly, and all the resellers have to do is find a new domain name that (checks notes) doesn't have the word DEEPSEEK in it. Like, seriously? Your goal was to identify resellers by checking if the proxy has the corporate name of one of your competitors in it? Is this amateur hour?<p>All Anthropic has done is reduce trust, once again, with legitimate customers, while doing nothing to stop illegitimate customers. They <i>need</i> to get adults into key leadership roles, quickly.</p>
]]></description><pubDate>Tue, 30 Jun 2026 17:00:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=48735676</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48735676</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48735676</guid></item><item><title><![CDATA[New comment by 827a in "Qwen 3.6 27B is the sweet spot for local development"]]></title><description><![CDATA[
<p>Apple does not sell a 64GB variant of the M4 Mac Mini. IIRC they never have; its always capped out at 48GB.<p>If you were planning on getting an M5 128GB; just get a DGX Spark (~$4500) or a 5090-equipped machine (~$4500) plus a Macbook Air (~$1500). You'll come in below the M5 Max 128 pricing (~$6700+ USD) and be happier for it.</p>
]]></description><pubDate>Tue, 30 Jun 2026 00:33:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48727145</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48727145</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48727145</guid></item><item><title><![CDATA[New comment by 827a in "AI's Affordability Crisis"]]></title><description><![CDATA[
<p>Yeah I just mean that if a business came to them and asked for fifty licenses to the $200/mo plan, OpenAI would tell them to kick dirt and basically pay API pricing. Startups should 100% just be telling their employees they can expense up-to $whatever/mo in AI-related expenses, and let software engineers go buy personal Codex/Claude subscriptions.</p>
]]></description><pubDate>Tue, 23 Jun 2026 22:53:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=48652649</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48652649</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48652649</guid></item><item><title><![CDATA[New comment by 827a in "AI's Affordability Crisis"]]></title><description><![CDATA[
<p>Company-wide their margins are trash (probably negative). They need as much inference margin as they can get to afford the massive training runs. It is likely that we'll see GPT-5.6 reduce API pricing to compete against Anthropic, but whether Anthropic feels they need to reduce their prices is anyone's guess.</p>
]]></description><pubDate>Tue, 23 Jun 2026 20:01:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48650516</link><dc:creator>827a</dc:creator><comments>https://news.ycombinator.com/item?id=48650516</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48650516</guid></item></channel></rss>