<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: rhdunn</title><link>https://news.ycombinator.com/user?id=rhdunn</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 07 Oct 2026 03:51:49 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=rhdunn" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by rhdunn in "Mistral Large 4"]]></title><description><![CDATA[
<p>There's also things like training costs, spend on AI data centres and GPUs, R&D, sales and marketing, etc. See e.g. <a href="https://www.reuters.com/business/finance/anthropics-ipo-prospectus-shows-sweeping-ai-vision-surging-costs-2026-09-28/" rel="nofollow">https://www.reuters.com/business/finance/anthropics-ipo-pros...</a>.</p>
]]></description><pubDate>Tue, 06 Oct 2026 17:55:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49981843</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49981843</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49981843</guid></item><item><title><![CDATA[New comment by rhdunn in "One month coding with GLM 5.3 Flash"]]></title><description><![CDATA[
<p>They are acting like they have infinite money (huge annual losses [1], [2], [3]) and infinite hardware (aiming to build a large number of data centres with a lot of GPU hardware [1], [3], [4]) even if that isn't the case.<p>[1] <a href="https://www.morningstar.com/stocks/anthropics-leaked-financials-reflect-fast-growth-not-2-trillion-valuation" rel="nofollow">https://www.morningstar.com/stocks/anthropics-leaked-financi...</a><p>[2] <a href="https://www.wheresyoured.at/anthropics-profitability-swindle/" rel="nofollow">https://www.wheresyoured.at/anthropics-profitability-swindle...</a><p>[3] <a href="https://fortune.com/2026/06/16/openai-financials-leaked-losses-revenue-profit/" rel="nofollow">https://fortune.com/2026/06/16/openai-financials-leaked-loss...</a><p>[4] <a href="https://www.forbes.com/sites/paulocarvao/2025/12/06/why-openais-ai-data-center-buildout-faces-a-2026-reality-check/" rel="nofollow">https://www.forbes.com/sites/paulocarvao/2025/12/06/why-open...</a></p>
]]></description><pubDate>Sun, 04 Oct 2026 07:01:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49951385</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49951385</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49951385</guid></item><item><title><![CDATA[New comment by rhdunn in "One month coding with GLM 5.3 Flash"]]></title><description><![CDATA[
<p>I suspect that some/a large portion of the build out is due to OpenAI/Anthropic/etc. just throwing hardware at the problem -- why bother trying to make training and inference efficient when you can just throw hardware/money at the problem. I remember NVIDIA talking about their supercomputer clusters with a large number of interconnected GPUs.<p>I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:<p>1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.<p>2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.<p>[1] Though the original idea for mixture of experts comes from a 1991 research paper (<a href="https://huggingface.co/blog/moe" rel="nofollow">https://huggingface.co/blog/moe</a>), so maybe a third area is research from Universities, etc.</p>
]]></description><pubDate>Sat, 03 Oct 2026 07:37:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49942161</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49942161</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49942161</guid></item><item><title><![CDATA[New comment by rhdunn in "What if we stopped using GPUs? [video]"]]></title><description><![CDATA[
<p>On the compute side there's the issue of scalability. CPUs are designed to perform a handful of operations at one time. They typically have a small number of dedicated integer, float, and other ALU configurations. Having dedicated matrix multiplication instructions would still lock that to how many matrix-capable ALUs there are in the CPU.<p>GPUs are designed to process a large number of calculations at once (so they can process triangles in 3D graphics). This makes them good at ML applications as they can process many of the matrix calculations at once. A 4090 has 16,384 CUDA cores (general compute ALUs) and 512 Tensor cores (dedicated matrix compute ALUs); a 5090 has 21,760 CUDA and 680 Tensor cores.<p>The other issue when training models (and running larger models) is the amount of VRAM (or RAM for CPUs) available. GPUs are limited in this aspect, whereas CPUs can have a lot higher memory. This affects things like batch size and the size of model that can be trained or fine-tuned.<p><a href="https://unsloth.ai">https://unsloth.ai</a> has guides for how to fine-tune existing models like Qwen 3.8 27B, memory requirements, etc.<p><a href="https://medium.com/@kailaspsudheer/the-transformers-arithmetic-527111099527" rel="nofollow">https://medium.com/@kailaspsudheer/the-transformers-arithmet...</a> has some information on training a base model. A 7B llama model is estimated at taking ~34GB memory for inference at F32, but was observed requiring 96GB memory when training (for the model weights, gradients, activations, and optimizer states).<p>Note: you can reduce the memory required for training by recomputing the gradients, at a cost of performance/time. You can also do other tricks like performing a QLoRA/LoRA pass on the model then merging that into the model to create a checkpoint.<p>I don't know what sized model you could train on 64GB/128GB RAM via a CPU.</p>
]]></description><pubDate>Fri, 02 Oct 2026 20:06:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49937978</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49937978</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49937978</guid></item><item><title><![CDATA[New comment by rhdunn in "Livenerf: Has Opus 5.5 been nerfed yet?"]]></title><description><![CDATA[
<p>I've seen that looping and glitching behaviour in over-quantized (~Q4 or lower) local models like Llama and Qwen 3.x. Thus, it is likely that they quantize the model after release to save on compute costs (while giving a favourable result at launch). That quantization can result in changes to the model's behaviour (you are changing the weights) that could be interpreted as nerfing.</p>
]]></description><pubDate>Wed, 30 Sep 2026 07:01:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49905396</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49905396</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49905396</guid></item><item><title><![CDATA[New comment by rhdunn in "When did Google get so weird?"]]></title><description><![CDATA[
<p>Google's search has been bad for a long time, esp. with older content. I had to use Bing to find a bug report from 2014 I was trying to find around 2 years ago.<p>It's a number of things like:<p>1. not searching for exact quotes/text, only words in the query;<p>2. including other lexical forms in the results (e.g. searching for things like "how do you open the run dialog" returns results for running (the sport/exercise) as well as run (an application));<p>3. including synonyms, typos, etc.<p>It seems like while each change makes sense and can improve results in some cases, over time the combined effect has made Google's search near unusable.</p>
]]></description><pubDate>Tue, 29 Sep 2026 06:33:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49889060</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49889060</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49889060</guid></item><item><title><![CDATA[New comment by rhdunn in "It's Time to Investigate the AI Labs"]]></title><description><![CDATA[
<p>What if you've been training your dog to bite people? (I.e. in OpenAI's case training their models around finding exploits in software, training on cyber security material, etc.)</p>
]]></description><pubDate>Tue, 29 Sep 2026 06:16:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49888943</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49888943</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49888943</guid></item><item><title><![CDATA[New comment by rhdunn in "Coding is not solved"]]></title><description><![CDATA[
<p>They matter when implementing the business logic of the application, when writing tests for a character counting function, etc.</p>
]]></description><pubDate>Mon, 28 Sep 2026 16:12:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49880337</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49880337</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49880337</guid></item><item><title><![CDATA[New comment by rhdunn in "Coding is not solved"]]></title><description><![CDATA[
<p>By not reviewing, reading, or understanding the code generated by agentic LLMs the output is effectively like a compiler. However, a compiler has deterministic behaviour that can be repeated and verified.<p>The behaviour/output of an LLM is not like that. Ask an LLM to create a dashboard to show games by genre and it will generate different results with each run, and each model/model version produces wildly different results.</p>
]]></description><pubDate>Mon, 28 Sep 2026 15:59:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49880119</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49880119</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49880119</guid></item><item><title><![CDATA[New comment by rhdunn in "U.S. appeals court upholds designation of Anthropic as supply chain risk"]]></title><description><![CDATA[
<p>It's going to make things complicated w.r.t. software used anywhere within and by the DoD:<p>1. the linux kernel has patches created by and security vulnerabilities identified by Claude/Anthropic;<p>2. same with other software like SQLite and rsync.</p>
]]></description><pubDate>Fri, 25 Sep 2026 20:19:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49849402</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49849402</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49849402</guid></item><item><title><![CDATA[New comment by rhdunn in "Gravity Seems Holographic. What Does That Mean for Reality?"]]></title><description><![CDATA[
<p>My understanding is that Hawking radiation is where:<p>1. a particle/anti-particle pair is created at the event horizon;<p>2. the particle is on the outside edge of the event horizon, so "escapes" the black hole;<p>3. the anti-particle is on the inside edge of the horizon, so decreases the size of the black hole due to particle/anti-particle annihilation (with a corresponding particle on the inside of the black hole).</p>
]]></description><pubDate>Fri, 25 Sep 2026 18:39:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49848309</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49848309</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49848309</guid></item><item><title><![CDATA[New comment by rhdunn in "Gemini 3.8 text-to-speech"]]></title><description><![CDATA[
<p>They have thousands of engineers, so it is very likely that different teams are working on each of those with the relevant domain knowledge.</p>
]]></description><pubDate>Thu, 24 Sep 2026 07:00:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49827181</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49827181</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49827181</guid></item><item><title><![CDATA[New comment by rhdunn in "Gemini 3.8 text-to-speech"]]></title><description><![CDATA[
<p>It depends on what you are after (quality, legibility, performance, etc.).<p>If you're after quality then Qwen3 TTS is a very good model esp. if you take some effort to craft a voice file. It is slow, so isn't practical for real-time voices (like assistants). It can also occasionally switch to a different voice to the one provided, so you may want to break up the text being processed.<p>I've not yet tried other recent/recentish models.<p>If you are after performance then two options from older models are:<p>1. flite with a HTS (Hidden Markov Model) voice like cmu_us_rms (male) or cmu_us_slt (female);<p>2. espeak/espeak-ng with an MBROLA (an Overlapped Add model) voice (mb-us1, mb-de5-en, etc.).<p>Alternatively, you could try using Qwen3 TTS or over voice changing model with the CMU Arctic (<a href="http://www.festvox.org/cmu_arctic/" rel="nofollow">http://www.festvox.org/cmu_arctic/</a>) voice data which includes audio for the rms and slt voices among others.<p>If you're feeling adventurous you could also try fine tuning one of the TTS models on that data to create a custom voice, though the data is likely to be in the training data for the voices, so using an audio sample may be sufficient depending on the TTS model.</p>
]]></description><pubDate>Wed, 23 Sep 2026 18:34:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49820470</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49820470</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49820470</guid></item><item><title><![CDATA[New comment by rhdunn in "No Easy Fix for Bogus Respondents in Online Opt-In Polls"]]></title><description><![CDATA[
<p>Fair, though it does show how things like that can skew polling results.</p>
]]></description><pubDate>Wed, 23 Sep 2026 11:03:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49814223</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49814223</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49814223</guid></item><item><title><![CDATA[New comment by rhdunn in "No Easy Fix for Bogus Respondents in Online Opt-In Polls"]]></title><description><![CDATA[
<p>Or Boaty McBoatface winning the name of a research ship (given to a submarine on the ship).</p>
]]></description><pubDate>Wed, 23 Sep 2026 08:52:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49813326</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49813326</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49813326</guid></item><item><title><![CDATA[New comment by rhdunn in "SAML: A fractal of bad design"]]></title><description><![CDATA[
<p>My goto when integrating SSO in a Java/JVM application via SAML, OAuth, JWT, and other authentication mechanisms is pac4j (<a href="https://github.com/pac4j/pac4j" rel="nofollow">https://github.com/pac4j/pac4j</a>). It has integration to and demos for various web frameworks like Spring and Scala Play. It does a lot of the heavy lifting for integrating SSO into an application.</p>
]]></description><pubDate>Wed, 23 Sep 2026 07:01:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49812608</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49812608</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49812608</guid></item><item><title><![CDATA[New comment by rhdunn in "OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005"]]></title><description><![CDATA[
<p>The two remaining Zodiac cyphers are very short: 13 characters and 32 characters respectively. As the messages don't share a cypher with the other messages they could theoretically be anything.<p>The YouTube channel <a href="https://www.youtube.com/@doranchak/videos" rel="nofollow">https://www.youtube.com/@doranchak/videos</a> by David Oranchak, one of the people who solved the Z340 cypher, has some more details on this as well as how the Z340 cypher was cracked.</p>
]]></description><pubDate>Tue, 22 Sep 2026 14:59:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49802454</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49802454</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49802454</guid></item><item><title><![CDATA[New comment by rhdunn in "Pirate Face Rescues LLM Models from Deletion"]]></title><description><![CDATA[
<p>So... distribute a LoRA (or equivalent) that modifies the base weights with the abliteration vectors. That makes sense as it would be possible to try different abliterations and keep the storage space down.</p>
]]></description><pubDate>Sun, 20 Sep 2026 17:41:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49778055</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49778055</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49778055</guid></item><item><title><![CDATA[New comment by rhdunn in "Why building a Rust LSP is hard"]]></title><description><![CDATA[
<p>I've not written a language server but have written a language plugin for IntelliJ.<p>I started with writing a correct recursive descent parser. I then extended it to detect, report, and recover from common syntax errors as I encountered them so that the parser is robust. And adding a parser test case for each of these (e.g. one test for each branch through an EBNF construction).<p>Some examples are:<p>1. missing keywords when the keyword can be detected from the current context (e.g. missing semicolon at the end of a statement);<p>2. using the wrong token (e.g. `:` instead of `::` in a C++ namespace qualified name);<p>3. detecting and ignoring whitespace in a whitespace-sensitive qualification (e.g. in XML QNames);<p>4. keeping in the prolog state (where functions are defined) when there are errors so that functions after the error don't get lost;<p>5. lexing incomplete literals like `10e` so they can be handled as integers in the parser and emitting an error for them.</p>
]]></description><pubDate>Sat, 19 Sep 2026 06:25:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49763907</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49763907</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49763907</guid></item><item><title><![CDATA[New comment by rhdunn in "Everybody's Lost Their Minds"]]></title><description><![CDATA[
<p>I find AI autocomplete and chat (asking specific questions, getting general code I can adapt, etc.) to be useful. I'm not sold on agents and agentic workflow.<p>When I tried using agents for something as a test it immediately started doing things instead of being part of a conversation. I want to be actively involved in the process, not sit back and let AI agents write/generate stuff that I'd have to/end up rewriting/modifying anyway when it goes against the design I have in mind.<p>The other thing I don't like about agents is their ability to run any command [1]. That seems like a nightmare w.r.t. the potential for leaking secrets (signing/access keys, etc.) or doing damage (deleting files, database tables, etc.).<p>[1] You can set the option to review every command it runs, but you're then just hand-holding the agent.</p>
]]></description><pubDate>Thu, 17 Sep 2026 21:53:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49747123</link><dc:creator>rhdunn</dc:creator><comments>https://news.ycombinator.com/item?id=49747123</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49747123</guid></item></channel></rss>