<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: entilzha</title><link>https://news.ycombinator.com/user?id=entilzha</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 11 Oct 2026 01:23:43 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=entilzha" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by entilzha in "Byte latent transformer: Patches scale better than tokens"]]></title><description><![CDATA[
<p>Great to see our paper here again! Since the paper release, we've also released model weights here for anyone interesting in building on top of it: <a href="https://huggingface.co/facebook/blt" rel="nofollow">https://huggingface.co/facebook/blt</a>. We also added HF Hub code to easily load the model <a href="https://github.com/facebookresearch/blt?tab=readme-ov-file#load-weights-via-hf-hub">https://github.com/facebookresearch/blt?tab=readme-ov-file#l...</a>.</p>
]]></description><pubDate>Mon, 12 May 2025 22:02:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=43967879</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=43967879</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43967879</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>At least I wasn't aware of this work, but thanks for the refs! I'm always curious to read papers from 10-20+ years ago that have similarly inspired ideas. If it makes sense, we'll mention those in the next related work update.</p>
]]></description><pubDate>Sat, 14 Dec 2024 20:05:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=42419097</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42419097</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42419097</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>I don't believe so, or at least if someone tried it didn't work well enough that I remember :). Some of the motivation for the architecture changes in encoding patches stemmed from finding FLOP efficient ways to express relationships between byte sequences. E.G., having a long context window makes sense when dealing with tokens, but you don't need as long as an attention window if you're attending byte sequences to make patch representations, since the patch representations will implicitly be part of a longer context window in terms of number of patches.</p>
]]></description><pubDate>Sat, 14 Dec 2024 20:01:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=42419088</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42419088</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42419088</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>(Author here)<p>If I understand your question right, this is one of the reasons BPE is nice and the parent liked it. For any character sequence, provided the characters are in the alphabet used to create the BPE vocab, there are no unknown words/sequences. One downside of some previous tokenization methods is you could have unknown/UNK tokens, EG dictionary based methods.<p>In our paper with bytes, we also avoid the UNK issue, since we can have an embedding for every possible byte, since it’s not that many (and for sequences of bytes we use hash embedding, although we did test n-gram lookups for the top K frequent byte n-grams in the training data).</p>
]]></description><pubDate>Sat, 14 Dec 2024 17:32:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=42418320</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42418320</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42418320</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>(Author Here)<p>Good description! Maybe what parent got mixed up on is an alternate way to view this is trying to chunk bytes to have roughly similar information. EG we initially tried a bunch of patching schemes, EG, keep a running total of entropy until the total exceeds a threshold, but ended up finding simple things worked better.<p>I’ll see if we can add more information about the small CNN in a next update to arXiv paper.</p>
]]></description><pubDate>Sat, 14 Dec 2024 15:12:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=42417582</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42417582</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42417582</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>(Author Here)<p>Related thought, I think BPE is quite a good, cheap inductive bias to have in a model, which is part of what made it challenging to scale better against. I also suspect this is part of why with less training FLOPs BPE is better (left side of figure 1), BLT has to expend some of its FLOPs budget to recover/learn some of this useful bias. With more training FLOPs this becomes a smaller fraction of the budget though leading to better scaling.</p>
]]></description><pubDate>Sat, 14 Dec 2024 15:06:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=42417550</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42417550</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42417550</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>(Author Here)<p>There is at least some work on character based modeling, but it hasn’t scaled well before. The challenge I think with something more adhoc for exceptional tokens is that it’s hard to see gains since they are by definition, infrequent. If the text is rare enough, BPE should produce many single byte tokens, so current models actually expend more compute on these rare sequences.<p>BLT scales well because it expends less compute (by patching) on more predictable (low entropy) byte sequences. Current models only to some degree get this benefit, if it’s a larger BPE token, but that only goes so far.<p>So it’s really two related, but different motivations.</p>
]]></description><pubDate>Sat, 14 Dec 2024 14:56:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=42417504</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42417504</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42417504</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>(Author Here)<p>In editing we couldn’t find a good place for this so cut it in the current version, but at one point had discussed a parallel with information density of speech as described by one paper. Essentially the paper found that in languages that were less information dense per syllable, speakers spoke faster to achieve similar information density as languages with higher density per syllable. You could see patching by entropy paralleling this if you consider that low entropy bytes in terms of Shannon entropy are less information dense.</p>
]]></description><pubDate>Sat, 14 Dec 2024 14:49:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=42417463</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42417463</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42417463</guid></item><item><title><![CDATA[New comment by entilzha in "Byte Latent Transformer: Patches Scale Better Than Tokens"]]></title><description><![CDATA[
<p>(Author Here)<p>Not sure what you mean by implicit? If you mean just treat bytes as tokens, one issue you run into is your sequence lengths get quite long, so compared to a regular token LLM, you can’t pack as many bytes in a batch, which means you’re pretty FLOP inefficient so scale worse. You could make the model smaller to compensate, but then the model isn’t as good.</p>
]]></description><pubDate>Sat, 14 Dec 2024 14:43:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=42417428</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42417428</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42417428</guid></item><item><title><![CDATA[New comment by entilzha in "Sharing new research, models, and datasets from Meta FAIR"]]></title><description><![CDATA[
<p>Author here :), I do think it’s a good direction to look into! That said, aside from it being a bit too much to do at once, you’d also have to be careful about how you distributed your FLOP budget across the hierarchy. With two levels, you can make one level (bytes/local encoder) FLOP efficient and the other (patches/global encoder) FLOP intensive. You’d also need to find a way to group patches into larger units. But ya, there are many directions to go from here!</p>
]]></description><pubDate>Sat, 14 Dec 2024 08:01:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=42415429</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=42415429</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42415429</guid></item><item><title><![CDATA[New comment by entilzha in "Show HN: Open-source obsidian.md sync server"]]></title><description><![CDATA[
<p>I tried a few a while back. What I really want is as close to 1-1 to obsidian UI as possible. I found with some of the plugins that it could be hit/miss on working correctly. If I were doing only markdown notes, then wouldn’t need obsidian ;)</p>
]]></description><pubDate>Thu, 24 Aug 2023 22:16:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=37255275</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=37255275</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=37255275</guid></item><item><title><![CDATA[New comment by entilzha in "Show HN: Open-source obsidian.md sync server"]]></title><description><![CDATA[
<p>While you’re here, a killer feature for me would be the ability to privately host obsidian sites (similar to publish). Even if it required subscribing to publish to download a tarball of the site (that isn’t public), it could still be worth it. My use case is sharing obsidian notes with non-users (eg coworkers) in a private way.</p>
]]></description><pubDate>Thu, 24 Aug 2023 22:07:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=37255197</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=37255197</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=37255197</guid></item><item><title><![CDATA[New comment by entilzha in "My productivity app for the past 12 years has been a single .txt file"]]></title><description><![CDATA[
<p>Any thoughts on how to access/modify on mobile without making it too cumbersome? I often think about todo on walk/train, but could see making it a computer only thing.</p>
]]></description><pubDate>Sat, 08 Feb 2020 20:15:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=22278063</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=22278063</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=22278063</guid></item><item><title><![CDATA[New comment by entilzha in "Why Some Doctors Purposely Misdiagnose Patients"]]></title><description><![CDATA[
<p>Totally agree on not trusting any one doctor. Nowadays, I basically assume doctors are narrow minded experts and do the broader thinking myself (by reading widely). It’s the old saying that when you have a hammer everything looks like a nail. I wish I realized the extent to which this is true earlier.<p>My issue happened to be with sesamoid/metatarsal fracture, flat rootedness, and poor gait. No doctor cared about my gait which turned out to be the root cause of my other symptoms. They all wanted to do surgery/orthotics and basically thought that would fix it alone vs combined with PT.</p>
]]></description><pubDate>Thu, 23 Jan 2020 08:47:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=22125937</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=22125937</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=22125937</guid></item><item><title><![CDATA[New comment by entilzha in "Ask HN: Keras on PyPy?"]]></title><description><![CDATA[
<p>Not entirely true. For good utilization you need both GPU/TPU ops to be fast (written in C), but that won’t get you far if your input pipeline (possibly written in python) is slow. I could imagine if all the TF calls work in PyPy, that it would help with throughout by speeding up the input pipeline and keeping the GPU/TPU saturated more effectively. One solution is to write everything to TFRecords, but for experimentation that’s kinda annoying.</p>
]]></description><pubDate>Fri, 03 Jan 2020 07:31:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=21943999</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=21943999</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=21943999</guid></item><item><title><![CDATA[New comment by entilzha in "Seeing How Computers Think Helps Humans Stump Machines and Reveal AI Weaknesses"]]></title><description><![CDATA[
<p>I’m a co-author and would be happy to answer questions about our work!</p>
]]></description><pubDate>Wed, 07 Aug 2019 22:12:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=20639716</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=20639716</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=20639716</guid></item><item><title><![CDATA[Seeing How Computers Think Helps Humans Stump Machines and Reveal AI Weaknesses]]></title><description><![CDATA[
<p>Article URL: <a href="https://cmns.umd.edu/news-events/features/4470">https://cmns.umd.edu/news-events/features/4470</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=20639707">https://news.ycombinator.com/item?id=20639707</a></p>
<p>Points: 1</p>
<p># Comments: 1</p>
]]></description><pubDate>Wed, 07 Aug 2019 22:11:35 +0000</pubDate><link>https://cmns.umd.edu/news-events/features/4470</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=20639707</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=20639707</guid></item><item><title><![CDATA[New comment by entilzha in "PyPy v7.0.0: triple release of 2.7, 3.5 and 3.6-alpha"]]></title><description><![CDATA[
<p>Anyone know if puppy works with deep learning libraries like pytorch/tensorflow, or if there are plans to do so? Not looking for numerical speed ups, but for speed ups in preprocessing code</p>
]]></description><pubDate>Mon, 11 Feb 2019 15:25:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=19135064</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=19135064</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=19135064</guid></item><item><title><![CDATA[New comment by entilzha in "Google Store warranty horror story: what to do next?"]]></title><description><![CDATA[
<p>A similar issue got me from recommending android and google services to friends/family to actively discouraging them (and migrating off of every service I could feasibly do). The WiFi chip on my nexus 6 burnt out and neither Motorola or google would replace it without paying several hundred dollars (motorola said it was a software issue and google said it was a hardware issue). Looking at forums this was a known issue across every nexus device.<p>I have zero patience for poor customer service.</p>
]]></description><pubDate>Fri, 24 Aug 2018 15:02:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=17835338</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=17835338</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=17835338</guid></item><item><title><![CDATA[New comment by entilzha in "Show HN: Python library for functional programming"]]></title><description><![CDATA[
<p>You could actually simplify even more with the trick used in a comment farther down (<a href="https://github.com/0101/pipetools" rel="nofollow">https://github.com/0101/pipetools</a>). That way you would implement __or__ on for example `pipe` and not have to wrap each function.<p>As it turns out, implementing that different syntax in pyfunctional wouldn't be too hard or API breaking I think. Mainly it would require 1) wrapping/exporting functions to a module you could bring into scope (eg `from functional import functions as F` or `from functional.functions import *`), 2) Writing wrapper code to provide something functionally similar to `pipe`.<p>On libraries, I swear I saw something a while back, but my googlefu just now didn't help me find it.</p>
]]></description><pubDate>Tue, 19 Dec 2017 04:23:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=15957741</link><dc:creator>entilzha</dc:creator><comments>https://news.ycombinator.com/item?id=15957741</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=15957741</guid></item></channel></rss>