<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: kgeist</title><link>https://news.ycombinator.com/user?id=kgeist</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Mon, 05 Oct 2026 00:18:06 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=kgeist" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by kgeist in "LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents"]]></title><description><![CDATA[
<p>Well, I have experience writing and deploying LLM inference engines, and seeing how the whole thing easily collapses when something goes slightly wrong doesn't instill confidence either that you can just hook it up to arbitrary tools and then expect it not to do random silly stuff ("emergent behavior," heh).<p>It's not only about some ML theory about RL or AI safety; just silly numerical bugs, caching bugs, etc. in the inference layer can already make it do unexpected "unaligned" things, and the whole thing is just hacks upon hacks to make a silly text autocomplete look somewhat semi-intelligent. Most "post-trained" models are pretty much as useless as base models without harnesses that do the heavy lifting. Have an extra space in the chat template and intelligence goes to zero - here's your "AI" :)</p>
]]></description><pubDate>Sun, 04 Oct 2026 20:55:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49957749</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49957749</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49957749</guid></item><item><title><![CDATA[New comment by kgeist in "LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents"]]></title><description><![CDATA[
<p>Of course, they learn to generate "tool calls" to achieve "goals" instead of random prose, but at the end of the day, it's still a text autocomplete engine masquerading as an AI. In the happy path, on a known task, the text generator generates a sequence of "tool calls" you expect it to generate, but move off the happy path slightly and all bets are off, there's a non-zero chance it will do something totally random you never expect, because at that point it just throws random stuff at the wall until it succeeds, thanks to brute force with pre-learned heuristics masquerading as intelligence (which is especially the case with "agent swarms," as in the HuggingFace incident).</p>
]]></description><pubDate>Sun, 04 Oct 2026 20:17:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49957400</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49957400</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49957400</guid></item><item><title><![CDATA[New comment by kgeist in "LeCun has "zero concerns" about AI wiping out humanity, recent "rogue" incidents"]]></title><description><![CDATA[
<p>"Emergent behavior" in this case really is, imho, "we didn't think through all the edge cases carefully enough". You know, a non-AI system can also accidentally wipe out all data or do some other real harm (see the Knight Capital's stock exchange bug) simply because the developers didn't catch the edge cases earlier, and no one calls that emergent behavior. It's just a buggy system.<p>"AI" systems can do greater harm because they are usually run in loops until they finish, and they are given "tools". A non-AI system could technically accomplish the same too, via sheer brute force/fuzzing, the advantage of LLMs is that they can take shortcuts and do it much faster, thanks to certain things already being in the training data, a sort of brute force with statistics-based heuristics.<p>LLMs at the core are just text autocomplete engines, and they literally have randomization applied during token selection to make outputs "more creative" so that models search for more unexpected solutions by trial and error (temperature > 0). Not to mention compression is lossy as well. So it's understandable from the start that the outputs of an LLM cannot be 100% stable and guaranteed. With this in mind, if a researcher takes this obviously unpredictable system and gives it tools without a well-thought sandbox, I don't see any difference in principle, from a developer writing  "if rand() == 13 { launch_nukes() } If someone wrote such a function, and it did launch nukes, no one would argue that the rand function is dangerous and will kill us all. The fault is in the author of the code who attaches dangerous tools to an obviously unstable/unpredictable system, doesn't think it through, and then cries "rand will kill us all" when something goes awry fully removing all responsibility from himself. It's not "AI" doing harm but people at OpenAI and Anthropic with their irresponsible behavior.</p>
]]></description><pubDate>Sun, 04 Oct 2026 12:28:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49953340</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49953340</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49953340</guid></item><item><title><![CDATA[New comment by kgeist in "Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia"]]></title><description><![CDATA[
<p>This seems to be a thin wrapper around libllama, what's the point? Llama.cpp already ships a web server. I don't see anything in the README that llama.cpp doesn't already support.</p>
]]></description><pubDate>Fri, 02 Oct 2026 16:03:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49935075</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49935075</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49935075</guid></item><item><title><![CDATA[New comment by kgeist in "Context Language Models"]]></title><description><![CDATA[
<p>>Do we have to reinvent MMUs for LLMs<p>MMUs are already emulated in engines like vLLM (paged attention).<p>>This allows the model to learn what is most important to maintain in context<p>DeepSeek's Lightning Fast Indexer already does something similar, although without context compaction. It identifies which tokens are most important to attend to, which allows the model to skip irrelevant ones. A similar idea could be used to remove unnecessary tokens from the context altogether, while somehow strengthening the representation of the important ones (increasing their attention weight, merging information from discarded tokens into them, or creating compressed summary representations)</p>
]]></description><pubDate>Fri, 02 Oct 2026 12:38:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49932813</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49932813</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49932813</guid></item><item><title><![CDATA[New comment by kgeist in "Mercury 2.5 LLM hits 770 tokens per second"]]></title><description><![CDATA[
<p>That's the problem with etching a model onto a chip: by the time you've designed the chip, manufactured it, tested it, shipped it, and deployed it, the model will be hopelessly outdated (with the current improvement rates). And when you want to update, you have to buy new chips instead of just uploading a new model file like now. When Taalas announced their chip, the model was already 1.5 years old (stone age by current standards). It's their first chip, so maybe they can streamline it, but the problem of having to update hardware every few months to keep up with the industry is not going anywhere.</p>
]]></description><pubDate>Thu, 24 Sep 2026 12:45:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49829846</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49829846</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49829846</guid></item><item><title><![CDATA[New comment by kgeist in "Pre-Greek: The lost language hidden within Ancient Greek"]]></title><description><![CDATA[
<p>>wine<p>It could be an original Proto-Indoeuropean word as well, because Greek has ὑιήν "grapewine" (< *wih₁-ēn), which follows PIE ablaut (weyh₁-ō ~ wih₁-ēn), which doesn't usually happen if it's just a borrowing of a foreign word. And the same root is found in Latin vitis "vine", Russian vit'sa "to twist (often about vines)" etc.</p>
]]></description><pubDate>Fri, 18 Sep 2026 09:18:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49751908</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49751908</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49751908</guid></item><item><title><![CDATA[New comment by kgeist in "Pre-Greek: The lost language hidden within Ancient Greek"]]></title><description><![CDATA[
<p>>Words borrowed from a Pre–Indo-European language into Mediterranean languages<p>> [...] Greek μύρμηξ mýrmēx ‘ant’, Latin formica<p>Must be an error:<p><pre><code>  Proto-Celtic *morwos
  Proto-Balto-Slavic: *marwis
  Proto-Indo-Iranian: *marwiš
  Proto-Germanic: *mauraz
  Old Armenian: mrǰimn
</code></pre>
Greek murmēx could be an assimilation murw- => murm-, and Latin had dissimilation morm- => form- (although not clear what came first, maybe morm- was the original and morw- came later). Sanskrit also has vamra "ant", which makes it look like the whole thing is a tabooistic distortion of *wr̥mis "worm".<p>In any way, it doesn't look like it must be borrowed. Historically, some words once labeled Pre-Indoeuropean turned out to have pretty mundane PIE origins.</p>
]]></description><pubDate>Fri, 18 Sep 2026 08:54:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49751755</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49751755</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49751755</guid></item><item><title><![CDATA[New comment by kgeist in "How GLM built its own inference infrastructure"]]></title><description><![CDATA[
<p>I have a similar approach where I optimize kernels and find numerical differences between the CPU oracle and CUDA kernels using an automated AI agent in a feedback loop. Usually it solves numerical problems easily (it compares outputs of every layer and finds where they diverge), but so far no matter how many different SOTA models I throw at it, and even show it reference code from other inference engines, they aren't able to much the speed (my engine has a modification which is not found in reference code, although a lot of stuff is similar). Either I'm doing something wrong, or z.ai's Infra Agent is actually an agent swarm, i.e. a bruteforce with heuristics. My project is 2 weeks old so maybe I just need more time.</p>
]]></description><pubDate>Thu, 17 Sep 2026 16:31:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49743184</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49743184</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49743184</guid></item><item><title><![CDATA[New comment by kgeist in "RTK reports token savings, but our cost benchmarks disagree"]]></title><description><![CDATA[
<p>Judging by the leaks, OpenAI and Anthropic already train reasoning traces to use fewer tokens (they deliberately omit articles and prepositions, use very short sentences, etc.), even though you pay per token. So it wouldn't make sense to do that if the only incentive was "make them pay for as many tokens as possible per task."<p>It's more subtle than that. If a user has to wait longer for a solution/pay more, they'll be less satisfied and may switch to a competitor. More unnecessary tokens also means more unnecessary compute. Longer sessions are increasingly more expensive to serve than shorter sessions.<p>And there's always the Jevons effect: as a resource becomes cheaper, demand often increases, and so does net resource consumption.<p>So, imho, frontier labs have every incentive to reduce token usage per task (while also making you use AI for more and more tasks in your daily life)</p>
]]></description><pubDate>Fri, 11 Sep 2026 21:17:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49665508</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49665508</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49665508</guid></item><item><title><![CDATA[New comment by kgeist in "RTK reports token savings, but our cost benchmarks disagree"]]></title><description><![CDATA[
<p>The technique looked dubious from the start, because LLMs were trained to expect certain outputs from common bash tools. If the output is not what it expects, an LLM may issue more tool calls than before, because it will assume the tool is broken, the arguments passed to it were wrong, or it's a newer/older version of the tool etc => more tokens. Sounds like just adding to the prompt to use `grep` and `tail` extensively will do the trick without any special tooling.</p>
]]></description><pubDate>Fri, 11 Sep 2026 13:13:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49657868</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49657868</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49657868</guid></item><item><title><![CDATA[New comment by kgeist in "GPT-6 Astra, looped transformers, and hidden reasoning"]]></title><description><![CDATA[
<p>>What stops them <..> simply use cheaper model for every Nth request.<p>That would trigger a full prefill (context recompute) every Nth request because cached tokens aren't interchangeable between models, and that would require way more compute than just staying on Astra.<p>To avoid full recompute, you could prefill a cheaper model's context incrementally by always feeding it Astra's outputs in the background (and vice versa), but then that would require 1.5-2 more VRAM for each session + the complexity of keeping them in sync.<p>If the rumors are true that Astra is a looped transformer, a more practical approach would be to dynamically adjust the loop count during peak hours.</p>
]]></description><pubDate>Thu, 10 Sep 2026 01:45:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49637299</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49637299</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49637299</guid></item><item><title><![CDATA[New comment by kgeist in "DeepSeek launching v4.1 flash cheaper and more capable than v4 pro"]]></title><description><![CDATA[
<p>The web UI's system prompt is also probably in Chinese</p>
]]></description><pubDate>Wed, 09 Sep 2026 13:53:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49626615</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49626615</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49626615</guid></item><item><title><![CDATA[New comment by kgeist in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>They mention it here: <a href="https://github.com/argonautlabsai/deltafin/blob/main/k3-public-bench/results/PREFILL.md" rel="nofollow">https://github.com/argonautlabsai/deltafin/blob/main/k3-publ...</a><p>>device bytes read during the prefill window, all four drives (arm csv) 8,977 GB at 24.1 GB/s aggregate<p>I.e. low memory bandwidth.</p>
]]></description><pubDate>Wed, 09 Sep 2026 04:20:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49620937</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49620937</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49620937</guid></item><item><title><![CDATA[New comment by kgeist in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>There's a tendency to cite only decode speeds, but in practice, an LLM generates far fewer tokens than it has to read (unless you ask general knowledge questions). So the effective performance is much slower than the decode rate suggests, because a 512-token prompt already takes 6 minutes to load</p>
]]></description><pubDate>Wed, 09 Sep 2026 04:05:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49620801</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49620801</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49620801</guid></item><item><title><![CDATA[New comment by kgeist in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Do those reports require Kimi K3 though? Qwen3.6+ could probably do the same in a few seconds with similar quality.</p>
]]></description><pubDate>Wed, 09 Sep 2026 03:36:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49620595</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49620595</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49620595</guid></item><item><title><![CDATA[New comment by kgeist in "OpenAI begins rolling out GPT-6 Astra"]]></title><description><![CDATA[
<p>On their Agentic Index, GPT-6 Astra (both max/xhigh) has the same result as Qwen3.8-27b. Weird.</p>
]]></description><pubDate>Thu, 03 Sep 2026 20:10:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49556137</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49556137</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49556137</guid></item><item><title><![CDATA[New comment by kgeist in "The efficient frontier of LLM inference"]]></title><description><![CDATA[
<p>I'm currently trying to write an inference engine that combines the benefits of llama.cpp (one binary deployment, good support for heterogenous non-datacenter compute, wide quantization support) with the benefits of vLLM/SGlang (things like proper paged attention for better VRAM utilization and high concurrency).<p>Datacenter hardware is expensive and there's shortage of it but llama.cpp is slow/unoptimized for concurrent use, while vLLM/SGLang easily crash on non-common setups (things like, if you do pipeline parallelism for RTX5090+RTX4090, they will randomly crash with RAM caching enabled or select wrong kernels because they usually assume that every rank is the same device type; they also don't support Q5-Q6).<p>For me what's most interesting is to optimize inference for lack of good datacenter hardware and how to optimize for it best. I've been running an AI server in the office, and so far I've find these techniques most important for concurrent use on cheap hardware: pipeline parallelism (to accomodate for PCie), RAM caching (to quickly restore contexts into VRAM), speculative decoding (including domain-specific ngrams, they already can speed up code generation considerably without the overhead of a draft model), good kernels highly optimized for a specific device, support for Q5-Q6 (almost as good as Q8), FP8 contexts (more context to fit), paged attention (for better VRAM utilization), prefix caching, continuous batching (this is the default everywhere).<p>So far the main bottlenecks have been llama.cpp's poor VRAM utilization for contexts (you either have fixed-size slots, or use unified KV cache where each request attends to attention from all other requests and then unnecessary portions of attention are masked out), and lack of decode/prefill segregation: when a request starts prefilling a long context, all decoding threads slow down to like 5 tok/sec. On the other hand, vLLM/SGLang feel superbuggy if you don't run them on some officially approved node like 8xH200</p>
]]></description><pubDate>Wed, 02 Sep 2026 07:45:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49533089</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49533089</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49533089</guid></item><item><title><![CDATA[New comment by kgeist in "Terence Tao explains 6 essential mathematical concepts [video]"]]></title><description><![CDATA[
<p>>but also of other people.<p>Yeah, there's this thing called the curse of knowledge. If an engineer has a deep understanding of something, it's not a given that they can explain it well. For them, the topic feels so simple, and they've done it so many times that they may have forgotten other people aren't as knowledgeable. They will throw terms around without explaining them, etc.</p>
]]></description><pubDate>Tue, 01 Sep 2026 08:57:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49519630</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49519630</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49519630</guid></item><item><title><![CDATA[New comment by kgeist in "Creepy Crawlies"]]></title><description><![CDATA[
<p>How about: "Type the seahorse emoji to solve the CAPTCHA" :) Something that triggers infinite loops in LLMs or trips the guardrails.</p>
]]></description><pubDate>Sun, 30 Aug 2026 19:56:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49502176</link><dc:creator>kgeist</dc:creator><comments>https://news.ycombinator.com/item?id=49502176</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49502176</guid></item></channel></rss>