<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Majromax</title><link>https://news.ycombinator.com/user?id=Majromax</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 23 Aug 2026 04:55:59 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Majromax" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by Majromax in "I'm becoming AI-blind"]]></title><description><![CDATA[
<p>> But it’s just as likely to make an output better.<p>No, for any particular output token the model's true logits are definitionally the 'best' that the model can achieve.<p>This is inherently probabilistic.  The model's top-1 guess is not guaranteed to be optimal, but it should be so a proportionate fraction of the time.  Same with the top-2, top-3, etc.<p>Watermarking necessarily alters the output distribution away from the model-set distribution, and that alteration is inherently 'worse' in expectation.<p>You can liken this to a weather forecast.  If there's a 25% chance of rain, the forecast should say so (or a 'sampled' deterministic forecast should predict rain 25% of the time).  If the forecast is 'watermarked' and predicts rain 27% of the time under identical circumstances, it's a worse forecast.<p>That being said, this is a case of hiding a message in a noisy channel.  Watermarking only needs to communicate one bit ('yes watermark'), so the effects can be arbitrarily small provided one is willing to tolerate an increase to the text size needed for reliable detection.</p>
]]></description><pubDate>Sat, 22 Aug 2026 04:29:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49396568</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49396568</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49396568</guid></item><item><title><![CDATA[New comment by Majromax in "If this is true, the hyperscalers are toast"]]></title><description><![CDATA[
<p>> Hyperscalers don't run computing at some multiple more efficient than on prem.<p>I'd disagree here.  I see two avenues for an efficiency multiple, albeit a single-digit multiple:<p>* Client aggregation allows a hyperscaler to average out demand spikes from uncorrelated clients, reducing the peak:average demand ratio and allowing better budgeting of compute.<p>* Dynamic batching allows typical requests to run in batches of more-than-1 and/or overlap, offering better internal compute utilization ratios (e.g. interleaving output and input streams).  The small limit of on-device LLMs will run with batch sizes of one with strong memory bandwidth bottlenecks.<p>For an example of these factors in action, see the API cost differential between batch, standard, and 'fast' processing.  OpenAI prices these tiers at a 1:2:4 ratio.</p>
]]></description><pubDate>Thu, 20 Aug 2026 13:34:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49374383</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49374383</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49374383</guid></item><item><title><![CDATA[New comment by Majromax in "OpenRouter is joining Stripe"]]></title><description><![CDATA[
<p>> As things settle down and commoditize, the value of switching on a dime diminishes as people lock into their favorite models<p>I can imagine just the opposite outcome from the same scenario: as people settle into their favorite but commoditized models, competition for marginal inference cost will take over.  A company like OpenRouter that promises the cheapest tokens by the minute becomes essential on the low-cost margin.<p>I think that OpenRouter and equivalents get pushed out of the market only if the froth calms down (as you posit) <i>and</i> winning models stay proprietary, perhaps with their own unique API surfaces.</p>
]]></description><pubDate>Wed, 19 Aug 2026 18:45:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49365571</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49365571</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49365571</guid></item><item><title><![CDATA[New comment by Majromax in "AI usage patterns in software teams"]]></title><description><![CDATA[
<p>You could say the same thing about compilers versus assemblers, high-level languages versus low-level ones, and services and libraries versus monolithic programs.<p>All other things being equal, increasing the speed of some part of the development process will increase the overall pace of development.  However, By Amdahl's law that increase will be sublinear, and that is why we should take "pull requests" as an imperfect metric.<p>We also don't get to pick the form that 'better technology products' take.  While we'd probably like to keep cost(/effort) and complexity constant and increase robustness and performance, the market equilibrium might be 'worse is better' and reward whiz-bang features and lower effort.</p>
]]></description><pubDate>Wed, 19 Aug 2026 14:25:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49362073</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49362073</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49362073</guid></item><item><title><![CDATA[New comment by Majromax in "AI usage patterns in software teams"]]></title><description><![CDATA[
<p>>  Is the ROI there to pay for the trillions in commitments that have been bet on that ROI? That looks like a clear no at this point.<p>That's only a potential crisis for those who have made concrete investments.<p>On the use side, the 'cost' of AI spans more than two orders of magnitude.  Looking at recent models (<6mo) with reasonable performance (intelligence index >= 45) on OpenRouter, the output cost ranges from $50/MTok (Fable) to $0.153/MTok (DeepSeek Flash 0731).<p>From the perspective of a <i>user</i> of LLM/agent assistance, there's very likely a range where the benefits outweigh the costs.<p>If the ROI for the <i>model developers</i> isn't there, then that just impairs the future trajectory of the field.  Current models are just bits that aren't going anywhere, and as long as they can be served (in inference) above their marginal cost they will continue to be so-delivered.</p>
]]></description><pubDate>Wed, 19 Aug 2026 14:19:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49361993</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49361993</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49361993</guid></item><item><title><![CDATA[New comment by Majromax in "Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing"]]></title><description><![CDATA[
<p>A reasonable guess about the algorithm is 'A Watermark for Large Language Models' (<a href="https://arxiv.org/abs/2301.10226" rel="nofollow">https://arxiv.org/abs/2301.10226</a>).  The idea is that each generated token (or bigram) seeds a strong PRNG that splits the vocabulary into a 'green' and 'red' set.  The sampler then tries to select a 'green' next-token for generation.<p>After-the-fact checking only needs the vocabulary splitter, which is independent of the LLM.  Over a sufficiently large text non-watermarked text would expect to use green and red tokens with the baseline probability, and that difference can easily become statistically significant over sufficiently long texts.<p>The basic algorithm has obvious knobs to tune, among them the initial ratio of red to green tokens and how hard the sampler tries to pick a green token.  These would balance fidelity to the original distribution against watermark detectability (minimum required content length for statistical power).</p>
]]></description><pubDate>Mon, 17 Aug 2026 13:11:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49330282</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49330282</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49330282</guid></item><item><title><![CDATA[New comment by Majromax in "Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing"]]></title><description><![CDATA[
<p>> But unless you know the prompt, you don't fully know which choice the model faces. Surely a lot of coding space is wasted compensating for that uncertainty.<p>When you only need to encode one bit, the signal to noise ratio can be very low.  If I try to write my own human words under the policy of "try a little bit to avoid the letter 'e' in every fifth word," then a sufficiently long text would still be 'watermarked' even if I only succeed in this dictum (e.g.) 10% more often than the baseline.</p>
]]></description><pubDate>Mon, 17 Aug 2026 13:04:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49330180</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49330180</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49330180</guid></item><item><title><![CDATA[New comment by Majromax in "Going Dark, and the era of law enforcement hacking"]]></title><description><![CDATA[
<p>> Wouldn't one AI or another detect this deliberate backdoor and report it, as it'll look just like any other security vulnerability, the only difference being the intention?<p>That's precisely the author's point: deliberate backdoors will be more adversary-exploitable than ever before, but the demand for such from law enforcement agencies is likely to ratchet upwards.</p>
]]></description><pubDate>Fri, 14 Aug 2026 22:21:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49305326</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49305326</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49305326</guid></item><item><title><![CDATA[New comment by Majromax in "Don't classify, hallucinate!"]]></title><description><![CDATA[
<p>> In the notebook, I compute a MiniLM embedding of every real Wayfair classification. I compute the embedding of the fake, hypothetical embedding from the LLM. I then dot product the fake embedding into the real ones to find the most similar. Producing: [the right answer]<p>Isn't this begging the question that the hallucinated classification will be more selective with respect to the real schema than the query itself?  What would the dot product of <E(search query), E(schema)> have given?<p>Even if that is too vague, smaller LLMs are capable rerankers; return the top N matching true categories and ask for a contextual ordering.</p>
]]></description><pubDate>Fri, 14 Aug 2026 13:44:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49298600</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49298600</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49298600</guid></item><item><title><![CDATA[New comment by Majromax in "GLM-5.3: Frontier coding with emergent cyber capabilities"]]></title><description><![CDATA[
<p>> [A]ll of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.<p>This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots.  Frontier models should large dominate their competitors on a capability basis, but if GLM 5.2 (now 5.3) is routinely finding bugs / vulnerabilities missed by Fable and Sol then GLM might be genuinely a frontier-grade model by itself.</p>
]]></description><pubDate>Fri, 14 Aug 2026 13:29:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49298422</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49298422</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49298422</guid></item><item><title><![CDATA[New comment by Majromax in "Show HN: DeepSeek-V4 Latent Reasoning – moving "thinking" into latent space"]]></title><description><![CDATA[
<p>In my view, it's not so much the writing style itself as the lack of 'taste'.  Text that clearly seems AI-written has a flat level of exuberance that's just <i>exhausting</i>, kind of like a written version of the 'loudness war'.<p>Without some kind of dynamic range, I find myself having to do a lot of work to infer what points are truly important versus what are at best interesting implementation details.</p>
]]></description><pubDate>Sun, 09 Aug 2026 13:26:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49231211</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49231211</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49231211</guid></item><item><title><![CDATA[New comment by Majromax in "Our position on open-weights models"]]></title><description><![CDATA[
<p>> The attacker just needs one exploit chain, whereas the defender needs to block every avenue. Open access to models with no guardrails greatly benefits the attackers more than the defenders.<p>I see it as the opposite, where the attacker needs to find an exploit chain whereas the defender can block any link.<p>In this model, the balance of convenience favours the defender.  The defender presumably has access to the source code and configuration, so their scope of action is much larger than the attacker that must find vulnerabilities in a particular configuration.<p>I think that the different views might relate to different prior assumptions.  If we assume that each layer is mostly secure but may have a small number of latent vulnerabilities, then it should be relatively easy to find and fix those to create a perfectly secure layer.  If instead we assume that each layer is mostly <i>insecure</i> but chaining vulnerabilities is time-consuming then the land favours better-resourced attackers.<p>> Or find a remote exploit in Tesla cars and make their autopilot go on murdering rampages. (that one is from a movie)<p>In the worst case, air gaps and fixed contracts for information handling cover that.  Like any other domain, a car can be remotely exploitable only when untrusted information can influence behaviour inside the secured region.  Unfortunately, the convenience of OTA updates and 'cars as tech' rewards velocity at the expense of defensive design.</p>
]]></description><pubDate>Tue, 28 Jul 2026 16:54:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49086732</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49086732</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49086732</guid></item><item><title><![CDATA[New comment by Majromax in "Our position on open-weights models"]]></title><description><![CDATA[
<p>> The saying that stuck with me was "defenders have to be right 100% of the time, while attackers only have to be right once".<p>> You are suggesting this isn't correct?<p>The intuition behind that is applicable only when correctness is stochastic.  If you need to be waved in by a security guard, then one fake mustache might be the difference between being granted or denied entry.  However, a keypad either works or it doesn't; entering the wrong PIN is guaranteed refusal.<p>The other breach of that intuition is defense in depth.  Secure systems don't generally rely on a single binary trusted/untrusted status; the classified building still locks its interior doors.  This is the part that has – in my view temporarily – changed most with frontier models, in that they are much more skilled at chaining together vulnerabilities than previous models (and much <i>faster</i> about it than human experts, even if potentially less skilled).  If a system has a latent (0-day) vulnerability 50% of the time, then 10 independent layers would imply a ≈ 1/1000 chance that a critical compromise is possible.<p>However, these independent layers don't currently happen in practice because it's easier to write insecure code than secure code.  With luck, modest discipline, and defensive use of frontier models I think that this gap will narrow with time, in much the same way that it would be plainly crazy to deploy root access via telnet today.</p>
]]></description><pubDate>Tue, 28 Jul 2026 16:46:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49086612</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49086612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49086612</guid></item><item><title><![CDATA[New comment by Majromax in "Our position on open-weights models"]]></title><description><![CDATA[
<p>> Dumb question. If "Mythos-class" models are such a problem, then... why not just let it fix everyone's code?<p>In the specific case of cybersecurity, this is a reasonable medium-term outcome.  IMO, the cybersecurity risk is akin to the spread of a disease among an 'immune-naive' group: we can suddenly deploy much stronger attack-finding tools against large, established codebases created with much weaker security designs.  The path from here to there will be rough, but it's still fundamentally easier to write secure code than it is to exploit vulnerabilities.  (It's just easier yet to write insecure code, giving our status quo problem.)<p>For other 'safety' matters, defense isn't so easy because the attack and target are so different.  An AI propaganda bot or catfisher 'attacks' slowly-evolving human culture; one that instructs on explosives or bioterrorism directly interacts with an accomplice and not a victim.  If you believe that knowledge on how to build a pipe-bomb must be restricted, then giving everyone access to Fable does not mitigate the risk.<p>The controversial limit of this attitude is recursive self improvement and an AI singularity with potentially destructive results.  Proponents of this view think that sufficiently powerful AI is risky in nearly unimaginable ways such that the <i>capability itself</i> is harmful.  This is part (but not all) of why Fable (originally?) degraded itself when apparently assisting with AI research.</p>
]]></description><pubDate>Tue, 28 Jul 2026 16:33:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49086421</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49086421</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49086421</guid></item><item><title><![CDATA[New comment by Majromax in "Our position on open-weights models"]]></title><description><![CDATA[
<p>> Anthropic will make the case that their models should be evaluated with the safety layer in front, because that is the only way the model is available whereas open weight models need to pass the same test just on the weights.<p>Worse than that: an open-weight but safe model can be 'abliterated' to remove safety refusals using fine-tuning procedures that require a couple of orders of magnitude less compute than the original pretraining.<p>The 'universal evaluation' criterion then has three outcomes:<p>* It could become a mandatory, regulatory oversight of _all_ model training capable of hosting frontier-scale models.  Since GPUs for LLM training are the same GPUs for other model training, effective mandate would require GPUs be government owned or controlled as if they were weapons of mass destruction.<p>* It could impose limits on release of capable open-weight models, requiring Kimi et al to prove that they <i>cannot</i> be made capable of abusive behaviours.<p>* It could be security theatre.<p>The AI-as-existential-risk argument points towards the first, the competition-protection argument points towards the second, and least-effort implementation would be the last.</p>
]]></description><pubDate>Tue, 28 Jul 2026 16:24:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49086298</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49086298</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49086298</guid></item><item><title><![CDATA[New comment by Majromax in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>> I may be crucified for asking this but: is there any proof that slop matters beyond our sensibilities as developers?<p>That's precisely what this benchmark tries to quantify.  Since the benchmark incrementally expands the scope of each problem, 'sloppy' code is code that is hard to later modify.<p>I know this firsthand: the dumbest coder I've ever worked with was 'myself six months ago'.  That jackass never keeps the documentation up to date and hard-codes things that ought to be exposed as configuration.<p>The SlopCodeBench is an important but early-stage probe in this direction.</p>
]]></description><pubDate>Tue, 28 Jul 2026 13:30:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49083586</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=49083586</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49083586</guid></item><item><title><![CDATA[New comment by Majromax in "The Economics of Recursive Self-Improvement [pdf]"]]></title><description><![CDATA[
<p>> because people have other ideas <a href="https://en.wikipedia.org/wiki/Technological_singularity" rel="nofollow">https://en.wikipedia.org/wiki/Technological_singularity</a><p>In a weak sense, singularities are common and should be expected every so often.  In a mathematical sense, a singularity is where a model 'blows up' and neglected terms become part of the dominant balance.<p>Industrialization hit a point of diminishing returns, but the industrial revolution was nonetheless a 'singularity,' where life afterwards was qualitatively unpredictable to people who lived before.  Likewise, agriculture was such a technological singularity to hunter-gatherer ancestors.<p>I could even make a decent argument that writing and literacy were such a singularity, making inconceivable social organizations routine.<p>In that weak sense I <i>expect</i> AI to be a singularity, recursive self improvement or no.  Life in 2050 may be completely unpredictable to someone who was taken out of time in the year 2000.<p>The remaining questions are speed and intensity, and both of these questions are related to RSI.  If RSI works, then the 'fast takeoff' visions become more plausible where society transforms over months to a few years – at least locally where the enabling technologies have diffused.  If not, it might take a couple of decades.</p>
]]></description><pubDate>Tue, 14 Jul 2026 16:15:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48909090</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=48909090</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48909090</guid></item><item><title><![CDATA[New comment by Majromax in "Codex starts encrypting sub-agent prompts"]]></title><description><![CDATA[
<p>No, you'd still care.  YOLO mode is about instantaneous permissions and access control, inspection of subagent prompts is about retrospective <i>quality</i> control.  If the main model is instructing subagents to do a subtly wrong thing, the overall process quality will degrade in ways that might be very hard to detect or fix without deep inspection of the middle stages.</p>
]]></description><pubDate>Tue, 14 Jul 2026 15:07:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=48908042</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=48908042</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48908042</guid></item><item><title><![CDATA[New comment by Majromax in "A 1969 camera operators' strike created Upstairs Downstairs multiverse"]]></title><description><![CDATA[
<p>If you're deliberately displaying the image in black and white, the colour pattern is interference that should be suppressed.<p>However, this practice was not universal, and archivists have now recreated colour copies (<a href="https://en.wikipedia.org/wiki/Colour_recovery" rel="nofollow">https://en.wikipedia.org/wiki/Colour_recovery</a>) of some shows where only black and white recordings survived by reconstructing the colour from these interference signals.</p>
]]></description><pubDate>Sat, 20 Jun 2026 12:30:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48608755</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=48608755</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48608755</guid></item><item><title><![CDATA[New comment by Majromax in "Claude Fable is relentlessly proactive"]]></title><description><![CDATA[
<p>> I haven't yet had an agent rm -rf files.<p>That happened to me once; I was running one of a few free-tier models in a pi-coding-agent session.  The bash tool there is stateless and always begins from the launch directory, but the agent assumed state and executed `rm -rf .` intending to remove a build directory.  Instead it removed the whole project tree, including session logs and notes.<p>This was mostly a matter of amusement for me since I was running the agent inside a bubblewrap sandbox <i>for that very reason</i>, and the project itself was not very important.</p>
]]></description><pubDate>Fri, 12 Jun 2026 13:35:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=48503902</link><dc:creator>Majromax</dc:creator><comments>https://news.ycombinator.com/item?id=48503902</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48503902</guid></item></channel></rss>