<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: nvtop</title><link>https://news.ycombinator.com/user?id=nvtop</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 10 Oct 2026 02:59:06 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=nvtop" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by nvtop in "Whistle: Speech to Text in 16.9 MB"]]></title><description><![CDATA[
<p>I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you.</p>
]]></description><pubDate>Thu, 08 Oct 2026 18:21:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=50009778</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=50009778</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50009778</guid></item><item><title><![CDATA[Is mathematics over, or just graduating?]]></title><description><![CDATA[
<p>Article URL: <a href="https://docs.google.com/document/d/e/2PACX-1vTSh-pyNP3Gi99WMmsinnLmE9V5CDI0HEm6WGbIPMNt3V5SGlClHF8-BetotKHOImrvQDSXmFdiOw8D/pub">https://docs.google.com/document/d/e/2PACX-1vTSh-pyNP3Gi99WMmsinnLmE9V5CDI0HEm6WGbIPMNt3V5SGlClHF8-BetotKHOImrvQDSXmFdiOw8D/pub</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49963382">https://news.ycombinator.com/item?id=49963382</a></p>
<p>Points: 40</p>
<p># Comments: 58</p>
]]></description><pubDate>Mon, 05 Oct 2026 11:16:32 +0000</pubDate><link>https://docs.google.com/document/d/e/2PACX-1vTSh-pyNP3Gi99WMmsinnLmE9V5CDI0HEm6WGbIPMNt3V5SGlClHF8-BetotKHOImrvQDSXmFdiOw8D/pub</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=49963382</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49963382</guid></item><item><title><![CDATA[New comment by nvtop in "GNU Midnight Commander"]]></title><description><![CDATA[
<p>I tried to love mc, but its ergonomics felt slightly off. Maybe it's just hard to rewire my Norton / Volkov Commander / FAR Manager muscle memory, I don't know<p>I ended up being on a Linux fork of Far Manager, which works beautifully: <a href="https://github.com/elfmz/far2l" rel="nofollow">https://github.com/elfmz/far2l</a></p>
]]></description><pubDate>Wed, 17 Sep 2025 10:02:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=45273850</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=45273850</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45273850</guid></item><item><title><![CDATA[New comment by nvtop in "Train a 70b language model at home (2024)"]]></title><description><![CDATA[
<p>March 2024</p>
]]></description><pubDate>Thu, 24 Jul 2025 20:41:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=44675831</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=44675831</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44675831</guid></item><item><title><![CDATA[New comment by nvtop in "Mercury: Ultra-fast language models based on diffusion"]]></title><description><![CDATA[
<p>This video has a live coding part which implements a masked diffusion generation process: <a href="https://www.youtube.com/watch?v=oot4O9wMohw" rel="nofollow">https://www.youtube.com/watch?v=oot4O9wMohw</a></p>
]]></description><pubDate>Mon, 07 Jul 2025 13:51:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=44490381</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=44490381</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44490381</guid></item><item><title><![CDATA[New comment by nvtop in "Gemini Diffusion"]]></title><description><![CDATA[
<p>You can absolutely do it, and I think it's a nice idea to try.</p>
]]></description><pubDate>Thu, 22 May 2025 18:19:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=44065036</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=44065036</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44065036</guid></item><item><title><![CDATA[New comment by nvtop in "Gemini Diffusion"]]></title><description><![CDATA[
<p>Correct, diffusion LMs can edit their intermediate predictions, so "final" tokens aren't necessarily final. This is an exciting property because it allows models to correct errors in what's generated so far -- something that GPT-like models can't.<p>This editing is based on the Transformer's encoder property to predict token probabilities for __every__ token in a sequence, not just for [MASK]s. So when you input a sentence of three tokens `[MASK] cat barks`, Transformer will generate a probability distribution over the vocabulary for each of the three tokens, for free.<p>Now you can come up with many ways of how to decide whether you want to edit token or keep it as is. In the simplest case, take a new token if its probability higher than the original by some margin. In our example, say model returns the probability of the token "cat" on the second position as p_2("cat") = 0.3, while p_2("dog") = 0.6. We may want to replace "cat" with dog, and use it in the subsequent iterations.<p>Actual heuristics are slightly more complicated, but the base idea is this.<p>P.S. In order to teach LM not to just copy input unmasked tokens but to try to find a better replacement, your training objective should include replacing some % of input tokens with some other random token. Now you have part of the input masked, and part of the input corrupted, so the model can't blindly assume that all input tokens are here to stay.</p>
]]></description><pubDate>Thu, 22 May 2025 11:36:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=44061013</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=44061013</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44061013</guid></item><item><title><![CDATA[New comment by nvtop in "Gemini Diffusion"]]></title><description><![CDATA[
<p>Despite the name, diffusion LMs have little to do with image diffusion and are much closer to BERT and old good masked language modeling. Recall how BERT is trained:<p>1. Take a full sentence ("the cat sat on the mat")
2. Replace 15% of tokens with a [MASK] token ("the cat [MASK] on [MASK] mat")
3. Make the Transformer predict tokens at masked positions. It does it in parallel, via a single inference step.<p>Now, diffusion LMs take this idea further. BERT can recover 15% of masked tokens ("noise"), but why stop here. Let's train a model to recover texts with 30%, 50%, 90%, 100% of masked tokens.<p>Once you've trained that, in order to generate something from scratch, you start by feeding the model all [MASK]s. It will generate you mostly gibberish, but you can take some tokens (let's say, 10%) at random positions and assume that these tokens are generated ("final"). Next, you run another iteration of inference, this time input having 90% of masks and 10% of "final" tokens. Again, you mark 10% of new tokens as final. Continue, and in 10 steps you'll have generated a whole sequence. This is a core idea behind diffusion language models.<p>Of course, there are some optimizations in the real world. If you need to generate a really long text (over 200 tokens), you'd better split it in chunks and fully generate the first chunk in parallel before moving to the next one. This semi-autoregressive generation is what Block Diffusion does.<p>You can be smart about how exactly you pick tokens you consider generated and what % exactly. At earlier stages, when it's mostly noise, you can take more, and on final stages you can do more iterations and take fewer tokens.<p>All in all, diffusion LMs are still iterative, but the number of steps is much lower than in autoregressive models. A nice thing is that you can choose how many steps are you going to make, trading quality for speed.<p>In the extreme, you can even generate just one leftmost masked token with a diffusion LM, effectively turning it into a traditional causal language model.</p>
]]></description><pubDate>Thu, 22 May 2025 07:38:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=44059646</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=44059646</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44059646</guid></item><item><title><![CDATA[New comment by nvtop in "Understanding Reasoning LLMs"]]></title><description><![CDATA[
<p>I'm also very skeptical of the significance of this "aha moment". Even if they didn't include chain-of-thoughts to the base model's training data (unlikely), there are still plenty of it on the modern Internet. OpenAI released 800k of reasoning steps which are publicly available, github repositories, examples in CoT papers... It's definitely not a novel concept for a model, that it somehow discovered by its own.</p>
]]></description><pubDate>Fri, 07 Feb 2025 10:26:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=42971165</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=42971165</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42971165</guid></item><item><title><![CDATA[New comment by nvtop in "Introducing Copilot+ PCs"]]></title><description><![CDATA[
<p>The whole point of NPU-enabled devices is to run models locally, so they your data never leaves your device. This is a huge privacy win.</p>
]]></description><pubDate>Mon, 20 May 2024 19:09:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=40419000</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=40419000</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=40419000</guid></item><item><title><![CDATA[New comment by nvtop in "Zellij: A terminal workspace with batteries included"]]></title><description><![CDATA[
<p>I use tmux when SSH'ing to remote boxes, but when working locally I find native terminal panes and tabs to be a better experience. Does tmux provide anything extra to what wezterm/kitty/iterm2 do?</p>
]]></description><pubDate>Mon, 05 Feb 2024 11:31:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=39260050</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=39260050</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39260050</guid></item><item><title><![CDATA[New comment by nvtop in "NLP Course – For You"]]></title><description><![CDATA[
<p>A lot. POS taggers used to be linear classifiers + features. In 2018 they switched to BERT and similar encoder-only models. In 2023, POS tagging is largely irrelevant, because it was used as a part of a larger pipeline, but now you can have everything end-to-end with better accuracy by fine-tuning a sufficienly large pretrained model (LLM or encoder-decoder like T5)</p>
]]></description><pubDate>Sat, 23 Dec 2023 21:53:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=38748618</link><dc:creator>nvtop</dc:creator><comments>https://news.ycombinator.com/item?id=38748618</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38748618</guid></item></channel></rss>