<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: vladf</title><link>https://news.ycombinator.com/user?id=vladf</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 05 Sep 2026 07:40:51 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=vladf" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by vladf in "NanoGPT Slowrun: Language Modeling with Limited Data, Infinite Compute"]]></title><description><![CDATA[
<p>That still looks like a “converge faster” paper.<p><a href="https://arxiv.org/abs/2006.10732" rel="nofollow">https://arxiv.org/abs/2006.10732</a><p>The above provides a nuanced theoretical view. GD inductive bias is probably better unless your model is misspecified</p>
]]></description><pubDate>Thu, 05 Mar 2026 02:33:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=47256769</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=47256769</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47256769</guid></item><item><title><![CDATA[New comment by vladf in "Show HN: AutoThink – Boosts local LLM performance with adaptive reasoning"]]></title><description><![CDATA[
<p>This is available for Flash</p>
]]></description><pubDate>Wed, 28 May 2025 13:16:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=44115702</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=44115702</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44115702</guid></item><item><title><![CDATA[New comment by vladf in "Stuff I Learned at Carta"]]></title><description><![CDATA[
<p>why</p>
]]></description><pubDate>Sat, 24 May 2025 14:04:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=44081200</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=44081200</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44081200</guid></item><item><title><![CDATA[New comment by vladf in "4o Image Generation"]]></title><description><![CDATA[
<p>That's pretty disappointing, it has been out for a while, and we still get top comments like (<a href="https://news.ycombinator.com/item?id=43475043">https://news.ycombinator.com/item?id=43475043</a>) where people clearly think native image generation capability is new. Where do you usually get your updates from for this kind of thing?</p>
]]></description><pubDate>Tue, 25 Mar 2025 23:20:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=43477143</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=43477143</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43477143</guid></item><item><title><![CDATA[New comment by vladf in "New York claims a small victory in 'forever war on rats'"]]></title><description><![CDATA[
<p>Yes perhaps one day cities like Tokyo will catch up</p>
]]></description><pubDate>Tue, 04 Feb 2025 01:17:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=42926110</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=42926110</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42926110</guid></item><item><title><![CDATA[New comment by vladf in "Git Bisect-Find"]]></title><description><![CDATA[
<p>Ah! Finally a real world use for egg drop (with eggs and floors both equal to num commits since init, but maybe fewer eggs for those less patient).</p>
]]></description><pubDate>Sat, 20 Apr 2024 02:05:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=40094004</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=40094004</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=40094004</guid></item><item><title><![CDATA[New comment by vladf in "Towards 1-bit Machine Learning Models"]]></title><description><![CDATA[
<p>> You still have a group-size of 64 in 4-bit fyi.<p>Results may vary :)<p>> Again, and I keep repeating this but it seems to be ignored every time: this is experimental work and it's still in progress. This story of small group-sizes on large models should not be an issue.<p>Apologies if something I said (or I guess did not say...) offended you! It's a hypothetical, and one IME is not so easy to achieve, but maybe you have different results. So I didn't want to comment on this, maybe it's possible (but LLMs don't scale up as easily in terms of quantization than other networks like image classifiers, in my experience).<p>> The extreme quant buys you potentially 70x more efficient matmul via binary/ternary operations.<p>To be clear, such hardware does not yet exist, and it's unclear if you really can have more efficient binary/ternary matmul if you need high-precision accumulators and more frequent broadcasting shiftss. It's again a complicated hardware question to answer if the sum total latency of doing many high-precision accumulations and many scales/shifts will be smaller (or, chip-area-wise, even feasible to implement), compared to a 4-bit baseline.</p>
]]></description><pubDate>Mon, 01 Apr 2024 14:51:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=39894738</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39894738</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39894738</guid></item><item><title><![CDATA[New comment by vladf in "Towards 1-bit Machine Learning Models"]]></title><description><![CDATA[
<p>If you’re willing to pay for the latency cost of per layer cpu fetching/offloading, I don’t see what extreme quant buys you.<p>You could just do a layer-by-layer fetching scheme with 4 bit weights.<p>For training too, just fetch each layer twice per step as needed for fwd/bwd.<p>And all for hbm cost equal to one layer’s worth</p>
]]></description><pubDate>Sun, 31 Mar 2024 18:46:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=39886892</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39886892</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39886892</guid></item><item><title><![CDATA[New comment by vladf in "Towards 1-bit Machine Learning Models"]]></title><description><![CDATA[
<p>I see, so we’re still fetching the metadata to gpu, and rescaling on gpu, just on-demand and discarding metadata when we’re done with that layer?<p>Why not do the same optimization for layer weights themselves?</p>
]]></description><pubDate>Sun, 31 Mar 2024 14:53:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=39884800</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39884800</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39884800</guid></item><item><title><![CDATA[New comment by vladf in "Towards 1-bit Machine Learning Models"]]></title><description><![CDATA[
<p>Thanks for the reply. I’m quite familiar with subchannel quant, but still feel like my questions did not get addressed.<p>1 Could you post the full memory use of the methods? E.g. you include quip metadata in its GB but not hqq metadata in its GB.<p>2 If you have to go to cpu to shift and scale, how did you get latency lower than pure on device? Was this bsz1? No speculative decoding?<p>3 how can lora absorb shifts with only increasing rank by 1 if you have a shift per group?</p>
]]></description><pubDate>Sat, 30 Mar 2024 15:42:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=39875915</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39875915</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39875915</guid></item><item><title><![CDATA[New comment by vladf in "Towards 1-bit Machine Learning Models"]]></title><description><![CDATA[
<p>Err, you are just restating what I’m saying, without addressing the concerns.<p>1 - is it fair to use ram in two places and report only one of them without any asterisk? (If you think this is fair-oh boy wait till you hear about my 0GB hbm use inference algorithm)<p>2 - i know how subchannel quantization works. Are they hitting those reported latency numbers with per layer cpu pingpong to rescale?</p>
]]></description><pubDate>Sat, 30 Mar 2024 03:30:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=39871609</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39871609</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39871609</guid></item><item><title><![CDATA[New comment by vladf in "Towards 1-bit Machine Learning Models"]]></title><description><![CDATA[
<p>Really strong binary results. So strong it was fishy. I hope someone can explain my confusion below.<p>> We compared the performance of the Llama2-7B model in three configurations: FP16 (full precision), HQQ (without fine-tuning), and HQQ+ (with adapter layers) using a group-size of 8.<p>Interesting, what is "group-size of 8"?<p>From their HQQ post (<a href="https://mobiusml.github.io/hqq_blog/" rel="nofollow">https://mobiusml.github.io/hqq_blog/</a>), it's the block size at which they add scales (presumably 16-bit) and shifts (in that post, it's 8-bit).<p>So for every 8 binary weights we have a 16-bit scale and 8-bit shift?<p>> Fine-tuning with Low-Rank Adapters<p>They say they inline the shift into the LoRA but how can you do this, block-wise, without increasing your LoRA rank by num-blocks (they claim to only use 1 additional rank)?<p>Then, the reported 7B sizes, in GB:<p>> 13.5 (fp16) 1.76 (HQQ 1-bit) 1.85 (HQQ+ 1-bit) 2.72 (quip# 2-bit)<p>those numbers would make sense if it was _actually_ 1 bit. But if you include the overhead of 16-bit scales (and why is the shift inlineable into lora? still unexplained) it'd be more like 3-bit.<p>From their HF page:<p>> This version offloads the meta-data to the CPU, so only the binary weights and the low-rank adapters are stored in the GPU memory.<p>Interesting, so we have to go back to CPU to rescale? Is this how they counted GB? This should have been clearly caveated in the table. I also am amazed they got latency lower than quip if they pingpong to CPU.</p>
]]></description><pubDate>Sat, 30 Mar 2024 01:07:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=39870848</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39870848</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39870848</guid></item><item><title><![CDATA[Distillation with Linear Models]]></title><description><![CDATA[
<p>Article URL: <a href="https://vladfeinberg.com/2024/02/04/distillation-walkthrough.html">https://vladfeinberg.com/2024/02/04/distillation-walkthrough.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=39591692">https://news.ycombinator.com/item?id=39591692</a></p>
<p>Points: 3</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 04 Mar 2024 15:38:14 +0000</pubDate><link>https://vladfeinberg.com/2024/02/04/distillation-walkthrough.html</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39591692</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39591692</guid></item><item><title><![CDATA[New comment by vladf in "Jim Keller criticizes Nvidia's CUDA, x86"]]></title><description><![CDATA[
<p>I’m not sure if you’re being facetious, but this is literary available for early access, but not ga yet. <a href="https://simonwillison.net/2024/Feb/21/gemini-pro-video/" rel="nofollow">https://simonwillison.net/2024/Feb/21/gemini-pro-video/</a></p>
]]></description><pubDate>Fri, 23 Feb 2024 17:24:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=39483493</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39483493</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39483493</guid></item><item><title><![CDATA[New comment by vladf in "'Baby Bust': Why Fewer Young People Expect to Become Parents (2013)"]]></title><description><![CDATA[
<p>Isn't this literally the case in the US? You list dependents on your tax form.</p>
]]></description><pubDate>Sun, 11 Feb 2024 17:00:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=39336340</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39336340</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39336340</guid></item><item><title><![CDATA[New comment by vladf in "Ask HN: Do companies hire principal and staff level engineers from job postings?"]]></title><description><![CDATA[
<p>What is your current level? Is it shown on your LinkedIn?</p>
]]></description><pubDate>Tue, 16 Jan 2024 05:22:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=39009798</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=39009798</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=39009798</guid></item><item><title><![CDATA[New comment by vladf in "Push ifs up and fors down"]]></title><description><![CDATA[
<p>And yet, a Rust Option (or really any option) can just be viewed as a list of one or zero elements. <a href="https://rust-unofficial.github.io/patterns/idioms/option-iter.html" rel="nofollow noreferrer">https://rust-unofficial.github.io/patterns/idioms/option-ite...</a><p>In fact, in Haskell, operating on an option conditionally has the exact same functor as a list: `map`.<p>So what am I to do, with an iterator? It's conflicting advice! An if is a for for an option.</p>
]]></description><pubDate>Thu, 16 Nov 2023 02:49:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=38285371</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=38285371</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=38285371</guid></item><item><title><![CDATA[Crinkle Crankle Optimization]]></title><description><![CDATA[
<p>Article URL: <a href="https://vladfeinberg.com/2023/10/08/serpentine-wall.html">https://vladfeinberg.com/2023/10/08/serpentine-wall.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=37821248">https://news.ycombinator.com/item?id=37821248</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 09 Oct 2023 15:08:58 +0000</pubDate><link>https://vladfeinberg.com/2023/10/08/serpentine-wall.html</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=37821248</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=37821248</guid></item><item><title><![CDATA[New comment by vladf in "What scientists must know about hardware to write fast code (2020)"]]></title><description><![CDATA[
<p>I ended up needing this so often for graph processing, and for values which might be inexact if using floating point, that I saved the formula in a blog post. <a href="https://vladfeinberg.com/2020/03/07/subset-isomorphism.html" rel="nofollow noreferrer">https://vladfeinberg.com/2020/03/07/subset-isomorphism.html</a><p>The formula can be "oblivious" to the final size of the matrix too, which is helpful if you're doing some sparse ML training on edges (e.g., GNNs).</p>
]]></description><pubDate>Tue, 03 Oct 2023 18:20:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=37755623</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=37755623</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=37755623</guid></item><item><title><![CDATA[New comment by vladf in "Hutter Prize for compressing human knowledge"]]></title><description><![CDATA[
<p>What</p>
]]></description><pubDate>Wed, 13 Sep 2023 22:47:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=37502695</link><dc:creator>vladf</dc:creator><comments>https://news.ycombinator.com/item?id=37502695</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=37502695</guid></item></channel></rss>