<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mmastrac</title><link>https://news.ycombinator.com/user?id=mmastrac</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 07 Aug 2026 07:21:41 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mmastrac" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mmastrac in "Inside vLLM: Anatomy of a High-Throughput LLM Inference System (2025)"]]></title><description><![CDATA[
<p>I've been working on a fresh, AI assisted port of DiffusionGemma from scratch and it takes a significant amount of time to deslop. I've spend a nonzero amount of time on refactoring and comment-vomit cleanup.<p><a href="https://github.com/mmastrac/diffgemma" rel="nofollow">https://github.com/mmastrac/diffgemma</a></p>
]]></description><pubDate>Fri, 07 Aug 2026 03:42:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49205684</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49205684</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49205684</guid></item><item><title><![CDATA[New comment by mmastrac in "Does Speaking to Agents Like Cavemen Save 65% of Tokens? We Test"]]></title><description><![CDATA[
<p>I suspect this means that there's just a ~10% inefficiency in token to information mapping.</p>
]]></description><pubDate>Fri, 31 Jul 2026 03:30:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49118679</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49118679</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49118679</guid></item><item><title><![CDATA[New comment by mmastrac in "Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac"]]></title><description><![CDATA[
<p>Which one? Nemotron Diffusion? It's impossible to say for sure, but I have a fairly deep library of metal kernels that _might_ cover some of the nvidia model's architecture.</p>
]]></description><pubDate>Thu, 30 Jul 2026 04:00:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49105999</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49105999</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49105999</guid></item><item><title><![CDATA[New comment by mmastrac in "Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac"]]></title><description><![CDATA[
<p>Awesome. I may need to finally bite the bullet and upgrade my macOS to test out the MPP approach you've taken.<p>I've got a number of tiled-load kernels, and a top-k attention kernel that you might find interesting.</p>
]]></description><pubDate>Wed, 29 Jul 2026 17:54:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49100766</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49100766</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49100766</guid></item><item><title><![CDATA[New comment by mmastrac in "Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac"]]></title><description><![CDATA[
<p>TBH, I think there's some truth to that. I spent _ages_ tuning the kernels to match the tested FLOP count of my M3's processor. I only have an M3 though and wasn't able to push int8 very far on it, but I think there's a chance that M5-class machines and higher might have more capability in this regard.<p>What I also learned is that MLX/vLLM is probably within ~20% or so of the absolute max perf on Mac. I found some improvements over what they were doing, but we're at the point where it's challenging to optimize without per-stepping kernels.<p>I found a few improvements over stock DiffusionGemma along the way, like using top-k attention, which drastically improves perf on my mac without sacrificing any of the benchmarks I was able to throw at it.<p>FWIW some of the issues with Gemma being slow on Mac are specific choices they've made in the architecture that make it challenging to make use various optimizations that have popped up recently. I think a Kimi K3-style network hybrid with the diffusion bits of DiffusionGemma could have some serious sway.<p>I think that diffusion still has an edge locally, but with some architecture tweaks and CPU improvements it would actually be a winner (ie: training the network for smaller token batch sizes or flexibility in attention heads, a less expensive attention mechanism, and others).</p>
]]></description><pubDate>Wed, 29 Jul 2026 17:47:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49100670</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49100670</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49100670</guid></item><item><title><![CDATA[New comment by mmastrac in "Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac"]]></title><description><![CDATA[
<p>I have a project that's almost ready to run DiffusionGemma as well. The two project might potentially work well together. I'm getting ~20tok/s on a 36GB M3 and there's strong possibility we might be able to crib faster kernels from each other.<p>Feel free to reach out.<p>(currently at <a href="https://github.com/mmastrac/diffgemma" rel="nofollow">https://github.com/mmastrac/diffgemma</a> but not in a releasable state yet)</p>
]]></description><pubDate>Wed, 29 Jul 2026 16:22:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49099518</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49099518</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49099518</guid></item><item><title><![CDATA[New comment by mmastrac in "GrapheneOS Defends Data-Wiping Function That Blocked US Border Search"]]></title><description><![CDATA[
<p>The founding fathers would be aghast at what America became. Hard to imagine modern America passing some of those constitutional amendments that older America passed wayyy back.</p>
]]></description><pubDate>Tue, 28 Jul 2026 15:44:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49085621</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=49085621</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49085621</guid></item><item><title><![CDATA[New comment by mmastrac in "Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU"]]></title><description><![CDATA[
<p>Is it just me or does this post not mention how much RAM they had? I would love to know - I have a dual-Xeon 1U screamer with 96GB of DDR4 RDIMM just sitting around...<p>FWIW I'm getting a hardware max of 20 tok/s (approx topping out the GPU's compute) on my custom local diffusiongemma port running on an M3.</p>
]]></description><pubDate>Wed, 15 Jul 2026 17:21:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48924168</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48924168</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48924168</guid></item><item><title><![CDATA[New comment by mmastrac in "Ask HN: Add flag for AI-generated articles"]]></title><description><![CDATA[
<p>There's a bot user that seems to crosspost every lobsters story here if it isn't already.</p>
]]></description><pubDate>Mon, 13 Jul 2026 03:47:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48887659</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48887659</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48887659</guid></item><item><title><![CDATA[New comment by mmastrac in "Show HN: Getting GLM 5.2 running on my slow computer"]]></title><description><![CDATA[
<p>This sort of thing is a lot of fun.<p>I've been going smaller.. I have a custom-quantized Rust port of DiffusionGemma (26B) that seems to perform better (in responses) than benchmarks seemed to indicate and reasonably fast for its model size. Works really well on a 36GB mac as well for both prefill and generation.<p>It's been interesting learning about the balance of factors for performant metal kernels on unified memory.<p>Should have a repo up on github in the next few weeks.</p>
]]></description><pubDate>Thu, 09 Jul 2026 23:31:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48853830</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48853830</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48853830</guid></item><item><title><![CDATA[New comment by mmastrac in "Anatomy of a Failed (Nation-State?) Attack"]]></title><description><![CDATA[
<p>Author here - if anyone has any contacts at Cloudflare to get the proxied domains (at least roadpay[.]cc) taken down, that would be great. I wasn't able to get an abuse report to stick. Ditto for the related LinkedIn profile and Twitter accounts.<p>The C2 IP (89.124.107.161) and malware-serving git repo (144.124.244.92) are both hosted on VDSINA in Russia, so not sure if there's anything to do there.</p>
]]></description><pubDate>Sat, 27 Jun 2026 13:06:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48697920</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48697920</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48697920</guid></item><item><title><![CDATA[New comment by mmastrac in "MicroVMs: Run isolated sandboxes with full lifecycle control"]]></title><description><![CDATA[
<p>They are long-lived if you're a mayfly.<p>But I think the point is that they should be cheap to set up, and because of the short life, never really contain anything except the potential to compute when needed, not important data.</p>
]]></description><pubDate>Fri, 26 Jun 2026 16:49:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48688834</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48688834</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48688834</guid></item><item><title><![CDATA[Anatomy of a Failed (Nation-State?) Attack]]></title><description><![CDATA[
<p>Article URL: <a href="https://grack.com/blog/2026/06/25/dissecting-a-failed-nation-state-attack/">https://grack.com/blog/2026/06/25/dissecting-a-failed-nation-state-attack/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48687139">https://news.ycombinator.com/item?id=48687139</a></p>
<p>Points: 5</p>
<p># Comments: 0</p>
]]></description><pubDate>Fri, 26 Jun 2026 14:36:13 +0000</pubDate><link>https://grack.com/blog/2026/06/25/dissecting-a-failed-nation-state-attack/</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48687139</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48687139</guid></item><item><title><![CDATA[New comment by mmastrac in "Om Malik has died"]]></title><description><![CDATA[
<p>Om Malik was the guy who had the biggest influence on the direction of my life, by far. It was through him I met Naval Ravikant in 2007, and then through Naval I met my co-founder that led to my startup exit in the '10s.<p>Luck surface area. I really owe so much to Om. I really can't imagine where I would be without that chance.</p>
]]></description><pubDate>Fri, 26 Jun 2026 05:19:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=48682600</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48682600</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48682600</guid></item><item><title><![CDATA[New comment by mmastrac in "Deno 2.9"]]></title><description><![CDATA[
<p>One of the most fascinating things to fall out of the AI apocalypse is seeing how abundant AI access amplifies the qualities of a company.<p>IMO, Deno has always been more methodical, more focused (maybe too focused?) on standards. But now the Deno team is on the right track: using Claude extensively to improve the node.js compat which was absolutely herculean if not impossible before AI. [+]<p>On the other hand, Bun has always played a bit fast and loose, chasing metrics at the cost of stability. Access to abundant AI has sent that project off the rails.<p>Disclaimer: Former Deno engineer - I'm obviously going to have some biases. All IMO of course, but if you ask me I'd still bet on Deno in the long term, and I personally still use it for any .ts projects.<p>[+] There might be a dozen people in the world that know how sensitive and subtle the timing and ordering in the JS event loop is and how meticulous just this single part needs to be for major node.js projects not to completely crap themselves.</p>
]]></description><pubDate>Thu, 25 Jun 2026 17:23:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48676584</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48676584</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48676584</guid></item><item><title><![CDATA[New comment by mmastrac in "USB Power Delivery: Plugging into the Benefits"]]></title><description><![CDATA[
<p>The consequences of plugging the wrong voltage or polarity into a barrel jack, most of which are compatible enough, is pretty bad. Ask me how I know.</p>
]]></description><pubDate>Mon, 15 Jun 2026 04:55:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48536779</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48536779</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48536779</guid></item><item><title><![CDATA[New comment by mmastrac in "There Is Life Before Main in Rust"]]></title><description><![CDATA[
<p>FWIW this post has both a "thanks" section for the human reviewers, and numerous footnotes linking to more authoritative sources.</p>
]]></description><pubDate>Fri, 12 Jun 2026 22:08:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48510013</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48510013</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48510013</guid></item><item><title><![CDATA[New comment by mmastrac in "WASI 0.3"]]></title><description><![CDATA[
<p>I'd love it if WASI modules could introspect their own custom sections (potentially even more introspection than that), but I've never been able to figure out a good way to do this. Seems like a fairly useful feature for a few use cases.</p>
]]></description><pubDate>Fri, 12 Jun 2026 17:42:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48507151</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48507151</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48507151</guid></item><item><title><![CDATA[New comment by mmastrac in "There Is Life Before Main in Rust"]]></title><description><![CDATA[
<p>Author here, happy to answer any questions. I've been working on building some higher-level abstractions on link sections (specifically, link-time optimized collections like maps (1) and sorted slices (2)) and wanted to share the hard-fought knowledge from the last couple of months.<p>There's a decent amount of knowledge around pre-main work in Rust, but I think this is one of the first attempts to walk through mutable link sections, which open up a pretty wide world of optimization, IMO. Even without mutability, I figured there isn't nearly enough documentation on these approaches out there.<p>(1) <a href="https://docs.rs/scattered-collect/0.20.0/scattered_collect/map/struct.ScatteredMap.html" rel="nofollow">https://docs.rs/scattered-collect/0.20.0/scattered_collect/m...</a><p>(2) <a href="https://docs.rs/scattered-collect/0.20.0/scattered_collect/sorted_referenced_slice/struct.ScatteredSortedReferencedSlice.html" rel="nofollow">https://docs.rs/scattered-collect/0.20.0/scattered_collect/s...</a></p>
]]></description><pubDate>Fri, 12 Jun 2026 17:26:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=48506914</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48506914</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48506914</guid></item><item><title><![CDATA[There Is Life Before Main in Rust]]></title><description><![CDATA[
<p>Article URL: <a href="https://grack.com/blog/2026/06/11/life-before-main/">https://grack.com/blog/2026/06/11/life-before-main/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48493512">https://news.ycombinator.com/item?id=48493512</a></p>
<p>Points: 93</p>
<p># Comments: 25</p>
]]></description><pubDate>Thu, 11 Jun 2026 17:32:35 +0000</pubDate><link>https://grack.com/blog/2026/06/11/life-before-main/</link><dc:creator>mmastrac</dc:creator><comments>https://news.ycombinator.com/item?id=48493512</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48493512</guid></item></channel></rss>