<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Remnant44</title><link>https://news.ycombinator.com/user?id=Remnant44</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Mon, 24 Aug 2026 09:16:46 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Remnant44" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by Remnant44 in "There is no “done”: Reflections on a completed Appalachian Trail thru-hike (2022)"]]></title><description><![CDATA[
<p>This was an unexpected find on HN but I found it very moving. Good writing lets you do something that is almost impossible otherwise; to get a glimpse inside someone else's mind and experience, the thing that is forever hidden from us unassisted.<p>I think it's all too easy to view other humans almost like NPCs, and most of the interactions we have in our day to day life reinforce this. You interact with the grocery checker, the waiter, even unfortunately sometimes your loved ones, and those interactions are 'on rails', following a script.<p>What a nice change of pace from AI-written linked-in articles.</p>
]]></description><pubDate>Mon, 10 Aug 2026 17:47:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49247168</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=49247168</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49247168</guid></item><item><title><![CDATA[New comment by Remnant44 in "Exact, parallel 2D Delaunay triangulation for int32 coordinates"]]></title><description><![CDATA[
<p>I haven't had a chance to dig into the repo at all, so that's excellent, thank you!</p>
]]></description><pubDate>Thu, 06 Aug 2026 08:19:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49193984</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=49193984</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49193984</guid></item><item><title><![CDATA[New comment by Remnant44 in "Exact, parallel 2D Delaunay triangulation for int32 coordinates"]]></title><description><![CDATA[
<p>Looks very promising - I've been looking for a good delaunay library that supports constrained delaunay.<p>It looks like their performance benchmark is including multithreading, which although a useful feature, makes performance comparisons more difficult - would love to see a baseline single threaded performance as well.</p>
]]></description><pubDate>Thu, 06 Aug 2026 06:44:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49193320</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=49193320</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49193320</guid></item><item><title><![CDATA[New comment by Remnant44 in "The worthlessness of Vitamin D is mildly exaggerated"]]></title><description><![CDATA[
<p>Do you have any sources for this? I have also heard of rabbit starvation many times over the years, and it has always been in the context of too taking in too little fat -- essential fatty acids -- as well, due to the extreme leanness of the meat.</p>
]]></description><pubDate>Wed, 24 Jun 2026 17:55:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48663429</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=48663429</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48663429</guid></item><item><title><![CDATA[New comment by Remnant44 in "Nvidia is proposing a beast of a CPU system for Windows PCs"]]></title><description><![CDATA[
<p>I agree with you, but also:<p>outside of anything else, amdahls law means that as the parallel performance grows, we become _more_ limited by the inherently serial code, and thus single core performance, not less.<p>Given that single core performance is "harder" (can't just throw more cores/sockets at the problem), it's also critically important.</p>
]]></description><pubDate>Sat, 06 Jun 2026 17:50:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48427279</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=48427279</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48427279</guid></item><item><title><![CDATA[New comment by Remnant44 in "C++26 Shipped a SIMD Library Nobody Asked For"]]></title><description><![CDATA[
<p>It's a little dramatic to say avx512 is dead versus 10 - rather, I would say that avx10 finalizes a universally available set of avx512 extensions. For AVX 10.1, there's essentially, no difference after Intel backed out of reducing the vector length.<p>For at least the next decade AVX 512 will be the high performance target, reaching all of the zen4/5/6 CPUs as well as whatever avx-10 enabled CPUs Intel producers.</p>
]]></description><pubDate>Sun, 17 May 2026 18:13:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48171568</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=48171568</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48171568</guid></item><item><title><![CDATA[New comment by Remnant44 in "The happiest I've ever been"]]></title><description><![CDATA[
<p>There's a whole lot of us out there. I don't know if there's still a future in the thing that I love, which is where all the malaise comes from.</p>
]]></description><pubDate>Sat, 28 Feb 2026 20:02:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=47199580</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=47199580</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47199580</guid></item><item><title><![CDATA[New comment by Remnant44 in "AMD64 Bit Matrix Multiply and Bit Reversal Instructions"]]></title><description><![CDATA[
<p>I love me some isa extension, I'd love to know what these are intended and useful for for though. 1 bit inference? I hear they could be useful in crypto as well, but that's out of my field.</p>
]]></description><pubDate>Sun, 01 Feb 2026 03:55:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=46843471</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=46843471</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46843471</guid></item><item><title><![CDATA[New comment by Remnant44 in "Tell HN: Bending Spoons laid off almost everybody at Vimeo yesterday"]]></title><description><![CDATA[
<p>While you may be correct in the sense that, in a public acquisition statement, people should be inferring enormous context and not taking anything said at face value.<p>It's simultaneously true that this is the farthest thing from effective, honest, and clear communication. Reading between the lines here is required precisely because we all know that any acquisition statements made are, at best heavily coded, if not completely just fluff.<p>You can recognize that and still get angry that it's par for the course for such things to be not just devoid of useful information, but often actively deceiving.</p>
]]></description><pubDate>Wed, 21 Jan 2026 22:06:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=46712268</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=46712268</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46712268</guid></item><item><title><![CDATA[New comment by Remnant44 in "AVX-512: First Impressions on Performance and Programmability"]]></title><description><![CDATA[
<p>Fort what it's worth, I had the exact same experience you did when I started writing SIMD code explicitly with intrinsics.<p>I avoided it for a long time because, well, it was so damn ugly and verbose to do simple things. However, in actual practice it's not nearly as painful as it looks, and you get used to it quickly.</p>
]]></description><pubDate>Mon, 19 Jan 2026 19:05:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=46683059</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=46683059</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46683059</guid></item><item><title><![CDATA[New comment by Remnant44 in "AVX-512: First Impressions on Performance and Programmability"]]></title><description><![CDATA[
<p>AVX doesn't require alignment of any memory operands, with the exception of the specific load aligned instruction. So you/the compiler are free to use the reg,mem form interchangibly with unaligned data.<p>The penalty on modern machines is an extra cycle of latency and, when crossing a cacheline, half the throughput (AVX512 always crosses a cacheline since they are cacheline sized!). These are pretty mild penalties given what you gain! So while it's true that peak L1 cache performance is gained when everything is aligned.. the blocker is elsewhere for most real code.</p>
]]></description><pubDate>Mon, 19 Jan 2026 18:46:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=46682856</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=46682856</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46682856</guid></item><item><title><![CDATA[New comment by Remnant44 in "AVX-512: First Impressions on Performance and Programmability"]]></title><description><![CDATA[
<p>There are many situations where your data is essentially _majority_ unaligned. Considerable effort by the hardware guys has gone into making that situation work well.<p>A great example would be a convolution-kernel style code - with AVX512 you are using 64 bytes at a time (a whole cacheline), and sampling a +- N element neighborhood around a pixel. By definition most of those reads will be unaligned!<p>A lot of other great use cases for SIMD don't let you dictate the buffer alignment. If the code is constrained by bandwidth over compute, I have found it to be worth doing a head/body/tail situation where you do one misaligned iteration before doing the bulk of the work in alignment, but honestly for that to be worth it you have to be working almost completely out of L1 cache which is rare... otherwise you're going to be slowed down to L2 or memory speed anyways, at which point the half rate penalty doesn't really matter.<p>The early SSE-style instructions often favored making two aligned reads and then extracting your sliding window from that, but there's just no point doing that on modern hardware - it will be slower.</p>
]]></description><pubDate>Mon, 19 Jan 2026 07:17:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=46675870</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=46675870</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46675870</guid></item><item><title><![CDATA[New comment by Remnant44 in "AVX-512: First Impressions on Performance and Programmability"]]></title><description><![CDATA[
<p>which honestly, shouldn't be neccessary today with avx512. There's essentially no reason to prefer the aligned load/store commands over the unaligned ones - if the actual pointer is unaligned it will function correctly at half the throughput, while if it_is_ aligned you will get the same performance as the aligned-only load.<p>No reason for the compiler to balk at vectorizing unaligned data these days.</p>
]]></description><pubDate>Mon, 19 Jan 2026 04:53:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=46675181</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=46675181</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46675181</guid></item><item><title><![CDATA[New comment by Remnant44 in "The state of SIMD in Rust in 2025"]]></title><description><![CDATA[
<p>In practical use for simd, various min/max operations. On Intel at least, they propagate nan or not based on operand order</p>
]]></description><pubDate>Thu, 06 Nov 2025 09:04:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=45833049</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45833049</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45833049</guid></item><item><title><![CDATA[New comment by Remnant44 in "Better SRGB to Greyscale Conversion"]]></title><description><![CDATA[
<p>I've run into this as well. Problem is that linear RGB is most definitely not a perceptually uniform space, so blending in it frequently does something different than you want. Use linear for physically based light and mixing, but if you are modeling an operation that is based on human perception it is going to be completely wrong.<p>The dark irony then, is that sRGB with its gamma curve applied, models luminance better (closer to human perception) for blending than linear does. If you can afford to do the blend in a perceptually uniform space like oklab, even better of course.</p>
]]></description><pubDate>Sun, 19 Oct 2025 23:04:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=45638826</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45638826</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45638826</guid></item><item><title><![CDATA[New comment by Remnant44 in "Apple M5 chip"]]></title><description><![CDATA[
<p>Essentially ever other use case for a computer.<p>Whether you're playing games, or editing videos, or doing 3D work, or trying to digest the latest bloated react mess on some website.. ;)</p>
]]></description><pubDate>Wed, 15 Oct 2025 18:33:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=45596675</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45596675</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45596675</guid></item><item><title><![CDATA[New comment by Remnant44 in "Daniel Kahneman opted for assisted suicide in Switzerland"]]></title><description><![CDATA[
<p>I've had just the smallest touch of this caring for my elderly parents, and you have my deep empathy. It's exhausting and really really hard.</p>
]]></description><pubDate>Sun, 12 Oct 2025 09:05:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=45556658</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45556658</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45556658</guid></item><item><title><![CDATA[New comment by Remnant44 in "Why we need SIMD"]]></title><description><![CDATA[
<p>totally - especially given how bandwidth constrained CPUs still are, going wider than 512 doesn't make much sense. 512 itself was a stretch for quite a long time (and all the negative press on the original implementations was a consequence of being not-quite-ready for primetime), but for current hardware I think it's perfect.<p>But 128bit is just ancient. If you're going to go to significant trouble to rewrite your code in SIMD, you want to at least get a decent perf return on investment!</p>
]]></description><pubDate>Wed, 08 Oct 2025 23:15:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=45521710</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45521710</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45521710</guid></item><item><title><![CDATA[New comment by Remnant44 in "Why we need SIMD"]]></title><description><![CDATA[
<p>Sure.. in detail and abstracted slightly, the byte table problem:<p>Maybe you're remapping RGB values [0..255] with a tone curve in graphics, or doing a mapping lookup of IDs to indexes in a set, or a permutation table, or .. well, there's a lot of use cases, right? This is essentially an arbitrary function lookup where the domain and range is on bytes.<p>It looks like this in scalar code:<p>transform_lut(byte* dest, const byte* src, int size, const byte* lut) {
  for (int i = 0; i < size; i++) {
    dest[i] = lut[src[i]];
  }
}<p>The function above is basically load/store limited - it's doing negligible arithmetic, just loading a byte from the source, using that to index a load into the table, and then storing the result to the destination. So two loads and a store per element. Zen5 has 4 load pipes and 2 store pipes, so our CPU can do two elements per cycle in scalar code. (Zen4 has only 1 store pipe, so 1 per cycle there)<p>Here's a snippet of the AVX512 version.<p>You load the lookup table into 4 registers outside the loop:<p><pre><code>  __m512i p0, p1, p2, p3;
  p0 = _mm512_load_epi8(lut);
  p1 = _mm512_load_epi8(lut + 64);
  p2 = _mm512_load_epi8(lut + 128);
  p3 = _mm512_load_epi8(lut + 192);
</code></pre>
Then, for each SIMD vector of 64 elements, use each lane's value as an index into the lookup table, just like the scalar version. Since we only can use 128 bytes, we DO have to do it twice, once for the lower and again for the upper half, and use a mask to choose between them appropriately on a per-element basis.<p><pre><code>  auto tLow  = _mm512_permutex2var_epi8(p0, x, p1);
  auto tHigh = _mm512_permutex2var_epi8(p2, x, p3);
</code></pre>
You can use _mm512_movepi8_mask to load the mask register. That instruction sets each lane is active if its high bit of the byte is set, which perfectly sets up our table. You could use the mask register directly on the second shuffle instruction or a later blend instruction, it doesn't really matter.<p>For every 64 bytes, the avx512 version has one load&store and does two permutes, which Zen5 can do at 2 a cycle. So 64 elements per cycle.<p>So our theoretical speedup here is ~32x over the scalar code! You could pull tricks like this with SSE and pshufb, but the size of the lookup table is too small to really be useful. Being able to do an arbitrary super-fast byte-byte transform is incredibly useful.</p>
]]></description><pubDate>Wed, 08 Oct 2025 22:53:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=45521529</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45521529</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45521529</guid></item><item><title><![CDATA[New comment by Remnant44 in "Why we need SIMD"]]></title><description><![CDATA[
<p>Yes and no. I think neon is undersized for today at 128bit registers -- if you're working with doubles for example, that's only two values per register, which is pretty anemic. Things like shuffles and other tricky bitops benefit from wider widths as well (see my other reply)</p>
]]></description><pubDate>Wed, 08 Oct 2025 19:59:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=45519965</link><dc:creator>Remnant44</dc:creator><comments>https://news.ycombinator.com/item?id=45519965</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45519965</guid></item></channel></rss>