<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: nodja</title><link>https://news.ycombinator.com/user?id=nodja</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 10 Oct 2026 23:19:47 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=nodja" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by nodja in "Don't be fooled–LLMs don't reason"]]></title><description><![CDATA[
<p>I don't know what you would call it, but reasoning/thinking/whatever it is, is a way for LLMs to tighten their sampling space, while allowing for wide sampling to still happen. This is akin to people brainstorming ideas.<p>I know that sounds confusing, let me break down how I think about this.<p>1. LLMs don't pick the token that ends up being used. This is by design, if the LLM gives a wide choice, it can better adapt to real world scenarios. i.e. generalize.<p>2. Without reasoning, this means that the LLM either locks in on whatever the sampler picked. Or decides mid-sentence/response to correct itself. This is what used to happen before reasoning, still happens if you turn reasoning off.<p>3. With reasoning, the LLM can make as many mistakes as it wants and explore its sampling space. Then use its vast pattern matching capabilities to decide which parts of the reasoning make sense and which were idiot ideas.<p>4. Enabling reasoning makes it so LLMs are much more confident on the final response, and the logits should theoretically all be near 99% on a single token for every token, i.e. much closer to greedy decoding. It analyzed all the possible options and figured out the best outcome, so a stray sample doesn't cause the answer to go awry.<p>This is why reasoning traces are filled with "but wait". I don't know if those were added in organically or artificially in the RL training, but regardless they're a good way to let the LLM keep generating other options and explore it's sampling space to the fullest.<p>Note: I haven't tested any of this and it's just my theory, but I'm sure if you really wanna know you can have claude run some smoke tests :)</p>
]]></description><pubDate>Fri, 02 Oct 2026 14:41:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49934066</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49934066</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49934066</guid></item><item><title><![CDATA[New comment by nodja in "PS5 Relapse Exploit"]]></title><description><![CDATA[
<p>PS5 emulation on x86 hardware will be more like running wine, so as long as your PC is as powerful as a PS5, performance should be close...eventually. The real performance bottleneck will be the GPU translation layer, much like converting DirectX to Vulkan isn't trivial, so isn't converting GNM (PS5's graphics API) to Vulkan, even tho the APIs are close (IIRC, they're cousin APIs)</p>
]]></description><pubDate>Tue, 29 Sep 2026 19:28:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49899053</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49899053</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49899053</guid></item><item><title><![CDATA[New comment by nodja in "Xiaomi Mimo 2.6 live post-training dashboard"]]></title><description><![CDATA[
<p>They exist to detect degradation. Datasets are not perfect and if a batch contains too much bad data it can ruin a run, also an opportunity to find bad data and improve the dataset filtering.</p>
]]></description><pubDate>Wed, 16 Sep 2026 22:19:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49733832</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49733832</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49733832</guid></item><item><title><![CDATA[New comment by nodja in "Nvidia is the central bank of AI"]]></title><description><![CDATA[
<p>Same reason they go into low margin with the PS4/PS5 and other projects. Market penetration.</p>
]]></description><pubDate>Sun, 13 Sep 2026 09:13:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49681766</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49681766</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49681766</guid></item><item><title><![CDATA[New comment by nodja in "LG says we're fake news [video]"]]></title><description><![CDATA[
<p>Youtube allows you to AB test with up to 3 titles and thumbnails.</p>
]]></description><pubDate>Sat, 12 Sep 2026 21:32:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49677431</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49677431</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49677431</guid></item><item><title><![CDATA[New comment by nodja in "Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU"]]></title><description><![CDATA[
<p>The whole site looks like and reads like AI slop. The outcomes also don't make any sense and don't feel rigorously tested (no, having claude test for you doesn't count as rigorous).</p>
]]></description><pubDate>Mon, 17 Aug 2026 15:38:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49332788</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49332788</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49332788</guid></item><item><title><![CDATA[New comment by nodja in "DeepSeek V4 Flash 0731"]]></title><description><![CDATA[
<p>I'm the same way, I have a very low/sporadic usage of any subscription I've tried. I now just use openrouter with DS4 pro/flash. It also gets rid of usage anxiety where I would try to justify the $20/month by forcing myself to use the tokens for projects as the weekly limit deadline neared.</p>
]]></description><pubDate>Fri, 07 Aug 2026 20:33:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49215882</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49215882</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49215882</guid></item><item><title><![CDATA[New comment by nodja in "Is AI reasoning right for the wrong reasons?"]]></title><description><![CDATA[
<p>The way I think about it is that it's unreasonable for a compute graph with a static number of operations to be able to answer both y=a*10 and something like y=((((x+x)*(x+1))/((2*x)+2))+((x*(x+3))/(x+3))-((x*x)/(x+1))+((x*x)/(x+1))-((x*(x+3))/(x+3))) in a single forward pass. Tokens are essentially a unit of work and can also be used for intermediate steps, not just final results.</p>
]]></description><pubDate>Fri, 31 Jul 2026 16:59:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49125755</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49125755</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49125755</guid></item><item><title><![CDATA[New comment by nodja in "DeepSeek V4 Flash 0731 Intelligence, Performance and Price Analysis"]]></title><description><![CDATA[
<p>I'm not a heavy user of agentic coding, but still use them quite a bit for some automation here and there. I've been going around shopping all the ~$10 subscriptions and I finally settled on openrouter + ds4 pro. The more intensive days cost me $1 and I set a $2 weekly limit which I've never hit the past 3 weeks, to me it's way cheaper than most subscriptions and I don't have to worry about maximizing my weekly quota/resets.</p>
]]></description><pubDate>Fri, 31 Jul 2026 16:08:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49124933</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49124933</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49124933</guid></item><item><title><![CDATA[New comment by nodja in "Gemini last models: temperature, top_p, and top_k are deprecated and ignored"]]></title><description><![CDATA[
<p>Yes, the whole list will always sum to 1 (100%) because there's lots of more sampling parameters. top_p, top_k and temperature are just the ones that affect output the most. Most parameters do math around assuming the list sums to 1 and order is not always the same, some software even lets you change the order around.<p>top_a is not very common and is better explained if I explain how the much more common min_p works. min_p filters out tokens below a certain threshold. The formula is <filter threshold> = <min_p> * <top token probability>. So if the top token has 0.5 probability, min_p = 0.1 would cut out tokens below 0.05. This is a tunable that lets you filter out other tokens depending on how confident the model is.<p>top_a is almost the same formula but you just square the <top token probability>. So <filter threshold> = <top_a> * <top token probability> ^ 2. This makes the filtering ramp up faster (cut out more tokens) if the model has a much more confident top choice, but keep more choices if the model is not so confident.</p>
]]></description><pubDate>Wed, 22 Jul 2026 10:01:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49004274</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49004274</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49004274</guid></item><item><title><![CDATA[New comment by nodja in "Gemini last models: temperature, top_p, and top_k are deprecated and ignored"]]></title><description><![CDATA[
<p>The posted answers are either behind a paywall or very obtuse so I'll just explain. I'll assume you know what tokens are.<p>A models output is not a single token, but a list with the probability for all the tokens that it knows, so we need to use a sampler to select the token that it's going to be the next token in the sentence. For example a simple greedy sampler will choose the token with the highest probability, but samplers normally pick a random token weighted by probability. A model usually knows about ~250 thousand tokens and the probability of some of these tokens are gonna be high, but the vast majority is close to but not actually 0% so there's a chance the sampler might pick some random token that doesn't make much sense, so we filter tokens.<p>top_k filters the tokens so that only the k top tokens are selected. So top_k=50 will filter those 250k tokens to only 50. This is assuming the list of tokens is sorted by probability.<p>top_p filters the top tokens until a percentage is accumulated. So if for example if you set the top_p to 0.6 and the model gave the top token a 0.5 (50%) probability and the second top token a 0.2, those 2 token accumulated to 0.7 which is greater than what you set it to (0.6) so no more tokens are selected. If this ran after top_k=50 it'll turn the list of 50 tokens into one of 2.<p>After each filter parameter is processed, the probability of the tokens is adjusted to sum to 1 (100%), Also note that order of operation here matters, i.e. top_p could be applied before top_k, but most providers follow what's on huggingface, I think I've only seen different implementation in certain local model hosting frameworks.</p>
]]></description><pubDate>Wed, 22 Jul 2026 05:55:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49002377</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=49002377</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49002377</guid></item><item><title><![CDATA[New comment by nodja in "LM Studio Bionic: the AI agent for open models"]]></title><description><![CDATA[
<p>Yup, it's the main reason I don't use LM studio more. I only use it to try out new models/quants, then use llama.cpp directly to host them. LM Studio also doesn't do stuff like audio input and often has bugs that pure llama.cpp doesn't so it can be a net negative for certain use cases.</p>
]]></description><pubDate>Thu, 16 Jul 2026 22:11:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48940941</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48940941</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48940941</guid></item><item><title><![CDATA[New comment by nodja in "Inkling: Our Open-Weights Model"]]></title><description><![CDATA[
<p>It doesn't matter until it does. If the chinese government decides that open weight model releases are no longer allowed, that's a lot of companies that can't release new models. Same with the US government, etc. Having diversity is important.</p>
]]></description><pubDate>Wed, 15 Jul 2026 22:28:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=48927982</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48927982</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48927982</guid></item><item><title><![CDATA[New comment by nodja in "Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor"]]></title><description><![CDATA[
<p>This hasn't been tested in court. But there's a high chance that model weights are not copyrightable, only the code to generate them is.<p>Cloud models are usually protected by trade secret laws, leaking them would get you in trouble. However if the model is made available publicly, as long as you don't break the law to get them, anything after that would be fair game unless Apple can prove that humans have significant authorship over the weights, which hasn't been tested and is a significant burden to prove/disprove.</p>
]]></description><pubDate>Mon, 13 Jul 2026 16:45:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48895299</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48895299</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48895299</guid></item><item><title><![CDATA[New comment by nodja in "I Changed My Name"]]></title><description><![CDATA[
<p>Yes. Having 4 names are quite common in Portugal, specially in certain areas. The names are usually structured like this: G1 G2 FM FF<p>G1 and G2 are given names. Usually 2 "first names" that you see in english, but there's common combos and sometimes there's a word joining them. Examples: "Maria Jesus" vs "Maria de Jesus". Some names are more common to be put first, but almost every name can be put in any order, example: "José António" vs "António José".<p>FM and FF are easy. FF is the family name of your father (your father's FF), and FM is the family name from your mother (your mother's FF).<p>Where I was raised 99% of my friends had 4 names structured like this, I only knew a few that didn't. When I moved to Lisbon the 3 name structure was much more common, dropping the second given name.<p>In Portugal there's rules for naming your kids (at least there were when I lived there), but I think in Brazil such rules don't exist. The author is brazillian but his name seems to follow the traditional portuguese naming style, as you guessed his name in english could be translated to "Robert Anthony Smith of Almeida" (Almeida is a portuguese town).</p>
]]></description><pubDate>Thu, 09 Jul 2026 22:14:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48853073</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48853073</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48853073</guid></item><item><title><![CDATA[New comment by nodja in "GLM 5.2 and the coming AI margin collapse"]]></title><description><![CDATA[
<p>> but I don't see any historical analogues.<p>The losers are quickly forgotten. Palm, Blackberry, AOL, MySpace. Yahoo, etc.<p>Software gets replaced all the time too, you even listed one and didn't realize. 15 years ago you'd call office irreplaceable, now you have to add gsuite to the mix, in 15 years there might be others. I know people that have never had office installed on their PC and use spreadsheets daily.<p>> It seems that enterprises will pay top dollar for service guarantees, integration, and someone they can sue.<p>Of course. But why pay $25 per million tokens for sonnet when you can pay $3 for GLM? Both probably running on AWS/Azure/Etc. under some third party.</p>
]]></description><pubDate>Tue, 07 Jul 2026 00:03:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48812086</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48812086</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48812086</guid></item><item><title><![CDATA[New comment by nodja in "Leaking YouTube creators' private videos"]]></title><description><![CDATA[
<p>Same here, first try I tried asking from the main studio page, and it didn't catch the comment at all despite being the latest comment.<p>When asking specifically from the video, it did fool the AI somewhat[1], but no link. I tried changing it to retrieve the revenue as that's probably a more sensitive/worthwhile metadata.<p>[1] <a href="https://i.imgur.com/YoDA8MJ.png" rel="nofollow">https://i.imgur.com/YoDA8MJ.png</a></p>
]]></description><pubDate>Sat, 04 Jul 2026 20:56:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48789015</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48789015</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48789015</guid></item><item><title><![CDATA[New comment by nodja in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>That's only correct for specific models and not what parent was referring to.<p>Stable Diffusion 3, an open weights model, was laughed at at release for not being able to even generate a woman laying in grass. The community attributed this to the heavy dataset filtering. Since then other open weights releases have been made with no NSFW capabilities and the community claims they're not as good as anatomy as well.<p>You can google "stable diffusion 3 woman in grass" and press the images tab to see how the model failed spectacularly.</p>
]]></description><pubDate>Wed, 01 Jul 2026 10:55:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=48744839</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48744839</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48744839</guid></item><item><title><![CDATA[New comment by nodja in "Ornith-1.0: self-improving open-source models for agentic coding"]]></title><description><![CDATA[
<p>Last thing you want a model to do is hallucinate a tool call and it's outputs...</p>
]]></description><pubDate>Mon, 29 Jun 2026 20:47:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48725010</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48725010</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48725010</guid></item><item><title><![CDATA[New comment by nodja in "GLM-5.2 – How to Run Locally"]]></title><description><![CDATA[
<p>Pipeline parallelism. Instead of splitting layers by row/column. You split at the layer edges. So instead of having this huge bottleneck of bandwidth you only need to transfer about 4KB per token when changing devices on a model like Qwen 3 30BA3.</p>
]]></description><pubDate>Tue, 23 Jun 2026 04:53:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48640479</link><dc:creator>nodja</dc:creator><comments>https://news.ycombinator.com/item?id=48640479</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48640479</guid></item></channel></rss>