<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Argonautlabs</title><link>https://news.ycombinator.com/user?id=Argonautlabs</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 18 Sep 2026 18:11:40 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Argonautlabs" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[Show HN: GLM-5.3 744B at 4 tok/s on a MacBook Pro, experts streamed from 4 SSDs]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/argonautlabsai/argodrive">https://github.com/argonautlabsai/argodrive</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49738954">https://news.ycombinator.com/item?id=49738954</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Thu, 17 Sep 2026 10:54:23 +0000</pubDate><link>https://github.com/argonautlabsai/argodrive</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49738954</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49738954</guid></item><item><title><![CDATA[New comment by Argonautlabs in "How GLM built its own inference infrastructure"]]></title><description><![CDATA[
<p>Different angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB 
by streaming the experts from NVMe SSDs instead of keeping them in memory.<p>One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2.<p>Method and numbers: <a href="https://github.com/argonautlabsai/argodrive" rel="nofollow">https://github.com/argonautlabsai/argodrive</a> (built on antirez/ds4).</p>
]]></description><pubDate>Thu, 17 Sep 2026 10:40:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49738842</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49738842</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49738842</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Deep Seek v4.1 M5 Max at 17 tokens/s"]]></title><description><![CDATA[
<p>We stream DeepSeek-V4.1-Flash (518 GB, 4-bit) from NVMe on a 128 GB M5 Max, because it doesn't fit in RAM.<p><pre><code>                        prompt processing      steady decode
    upstream, internal       16.23                 10.59
    ours, internal only      28.04  1.73x          14.38  1.36x
    ours, +1 external        36.88  2.27x          16.05  1.52x
    ours, +2 external        43.62  2.69x          17.38  1.64x

  The first fork row is the one that matters: same single internal SSD, no replicas, no enclosures.</code></pre></p>
]]></description><pubDate>Wed, 16 Sep 2026 04:58:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49722275</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49722275</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49722275</guid></item><item><title><![CDATA[Deep Seek v4.1 M5 Max at 17 tokens/s]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/argonautlabsai/argodrive">https://github.com/argonautlabsai/argodrive</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49722242">https://news.ycombinator.com/item?id=49722242</a></p>
<p>Points: 13</p>
<p># Comments: 2</p>
]]></description><pubDate>Wed, 16 Sep 2026 04:53:17 +0000</pubDate><link>https://github.com/argonautlabsai/argodrive</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49722242</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49722242</guid></item><item><title><![CDATA[New comment by Argonautlabs in "DeepSeek v4.1 Flash (518GB, 4-bit) on a 128GB MacBook: 2.7x prefill, 17 tok/s"]]></title><description><![CDATA[
<p>44 t/s prefill and 17 t/s decode on M5 Max 128gb Deep Seek V4.1 Flash (4 bit)</p>
]]></description><pubDate>Tue, 15 Sep 2026 13:35:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49712330</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49712330</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49712330</guid></item><item><title><![CDATA[DeepSeek v4.1 Flash (518GB, 4-bit) on a 128GB MacBook: 2.7x prefill, 17 tok/s]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/argonautlabsai/argodrive">https://github.com/argonautlabsai/argodrive</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49711198">https://news.ycombinator.com/item?id=49711198</a></p>
<p>Points: 1</p>
<p># Comments: 1</p>
]]></description><pubDate>Tue, 15 Sep 2026 11:57:25 +0000</pubDate><link>https://github.com/argonautlabsai/argodrive</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49711198</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49711198</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Often Deep Seek V4 flash or Qwen should be enough.<p>I wanted to see whether Kimi runs at all on one machine with the full record published, and for long multi-table finance reasoning I wanted the strongest model I could keep on the machine.<p>I did some tests against Deep Seek v4 flash results on my reports and Kimi definitely has some advantages.</p>
]]></description><pubDate>Wed, 09 Sep 2026 19:08:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49632380</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49632380</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49632380</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>OWC Express 1M2 (Thunderbolt 5, single M.2 NVMe) — three of them, two on the Mac's own ports and one behind an OWC Thunderbolt 5 hub since the machine has three ports.<p>Each enclosure tops out at about 7.1 GB/s on whole-file reads regardless of the drive inside (a 2 TB SN8100 measures the same as the 1 TB);<p>the drive behind the hub reads 5.7 GB/s and falls with queue depth.<p>Details in the README's hardware section</p>
]]></description><pubDate>Tue, 08 Sep 2026 23:04:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49618395</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49618395</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49618395</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>In principle yes, and the upstream engine already has a CUDA path with expert streaming and residency (that's theirs, not ours — we only measured on this Mac).<p>Two things carry over: the experts are read from disk per token either way, and the barrier model — a layer waits for the slowest of its 16 reads — is platform-independent.<p>Two things don't: the 50 GB resident trunk lives in unified memory here, so on a discrete GPU it would need to fit in VRAM or be streamed too; and a desktop's PCIe lanes let you put NVMe drives on the bus directly rather than behind a ~7 GB/s Thunderbolt enclosure, which is our per-drive wall.<p>Whether that ends up faster is exactly the kind of thing that wants measuring rather than guessing.</p>
]]></description><pubDate>Tue, 08 Sep 2026 22:55:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49618314</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49618314</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49618314</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Just internal 2Tb Macbook M5 Max drive it came in at roughly half the four-drive speed<p>(0.535 vs 1.038 tok/s at 128 tokens), since one fast drive still has to serve all 16 reads per layer while four drives split the load<p><a href="https://raw.githubusercontent.com/argonautlabsai/deltafin/main/k3-public-bench/results/charts/drive-draw.svg" rel="nofollow">https://raw.githubusercontent.com/argonautlabsai/deltafin/ma...</a></p>
]]></description><pubDate>Tue, 08 Sep 2026 22:20:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617993</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617993</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617993</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Memory, not the model.<p>The KV cache on this engine grows about 2.8 MiB per token of context, and the machine's 128 GB is already holding the 50.7 GiB resident trunk plus reserves the speculative-verify path needs<p>(we found the hard way that squeezing those makes the verifier reject wide batches and decode falls to single-token steps).<p>With the current reserves the engine admits ~4.4k tokens; that's a configuration ceiling you can raise by giving the cache more of the 128 GB and accepting less headroom elsewhere.<p>K3 itself supports far longer contexts — but see the prefill caveat above: on this setup long prompts cost minutes per 512 tokens until the scheduling fix lands.</p>
]]></description><pubDate>Tue, 08 Sep 2026 22:09:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617870</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617870</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617870</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Thank you.<p>Put a five-line TL;DR at the top of the README — what it is, the number, the honest limit, the two findings, credits — with the detail below for anyone who wants it.</p>
]]></description><pubDate>Tue, 08 Sep 2026 22:06:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617821</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617821</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617821</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Fair, and we didn't measure it.<p>Decode was flat from 128 to 512 generated tokens (0.926 → 0.923 tok/s drafter-off), but that's a 6-token prompt plus the output — total context under a thousand.<p>The current configuration admits about 4.4k tokens of context at all, and at anything like 200k the killer wouldn't be decode, it would be prefill: today it reads each layer's experts once per 64-row pass,<p>so 200k tokens of prompt would be measured in days, not minutes, until the scheduling fix.</p>
]]></description><pubDate>Tue, 08 Sep 2026 21:45:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617542</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617542</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617542</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>That's the one workload this setup is worst at today, unfortunately: output tokens are cheap at 1/s but input isn't — a 512-token prompt takes ~6 minutes before the first token,<p>because prefill currently reads each layer's experts once per 64-row pass (~9 TB of reads for a 1.4 TB model).<p>Fix is scheduling and it's the next thing being built; once prefill reads each expert once per layer, the one-token-out classifier pattern becomes the sweet spot rather than the worst case.</p>
]]></description><pubDate>Tue, 08 Sep 2026 21:41:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617497</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617497</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617497</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Probably not. Each read here is a whole 17.5 MB expert file, so the time per read is set by the drive's throughput, not its access latency — 17.5 MB at 7 GB/s is ~2.5 ms, which is what we measure at queue depth 1 on the SN8100s.<p>What actually moves the barrier is how fast the slowest of 16 concurrent whole-file reads completes: more drives on direct ports, stable tail behaviour under load, and scheduling. Our ladder shows even that with diminishing returns (one drive ≈52% of four, three ≈90%).</p>
]]></description><pubDate>Tue, 08 Sep 2026 21:38:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617459</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617459</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617459</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Bandwidth doesn't multiply like that here, and we measured it rather than assumed it. A MoE layer needs 16 expert reads and can't proceed until the slowest one lands, so a layer costs the max over its reads, not the sum.<p>Going from one drive to four (13.6 → ~33 GB/s of combined ceilings) took decode from ~52% to 100% of our number — not 4× — with<p>Every drive already at 90–100% of its own ceiling. RAID-0 was one of the first things tried and it lost: striping makes every read touch every drive, so the slowest drive sets every barrier.<p>What moves this is per-read latency and read scheduling, and for long prompts not re-reading each layer's experts eight times.<p>Numbers in results/SCALING.md and results/PREFILL.md.</p>
]]></description><pubDate>Tue, 08 Sep 2026 21:33:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617397</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617397</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617397</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>SSDs are connected via Thunderbolt 5 enclosures. I have one Gen4 and three Gen5 ssds inside enclosures. You can see specs here<p><a href="https://github.com/argonautlabsai/deltafin/tree/main/k3-public-bench" rel="nofollow">https://github.com/argonautlabsai/deltafin/tree/main/k3-publ...</a></p>
]]></description><pubDate>Tue, 08 Sep 2026 21:09:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617091</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617091</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617091</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Thank you! Here is the short version:<p><pre><code>  Kimi K3, 2.78T parameters, ~1.45 TB of MXFP4 experts streamed from four SSDs on an M5 Max / 128 GB. 1.00 tok/s steady over 512 tokens, 1.13 over 128, ~6.3 min to first token on a 512-token prompt. Output token-identical drafter on/off on a given drive layout; the int8 trunk is non-weight-exact per upstream.
  
  The useful bits: one drive gives ≈52% of four, two ≈73%, three ≈90%; and prefill is slow because of ~9 TB of reads for a 1.4 TB model — a scheduling bug with a planned fix.
  
  README with per-run logs: github.com/argonautlabsai/deltafin — a fork of gavamedia/deltafin, who built the engine.</code></pre></p>
]]></description><pubDate>Tue, 08 Sep 2026 21:04:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49617014</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49617014</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49617014</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>It actully does the job. Example: every morning it takes 30-40 minutes to generate reports automatically and these reports are being sent as a pdf to read to Telegram.</p>
]]></description><pubDate>Tue, 08 Sep 2026 20:46:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49616773</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49616773</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49616773</guid></item><item><title><![CDATA[New comment by Argonautlabs in "Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs"]]></title><description><![CDATA[
<p>Not useful for chat, agreed — and I wouldn't pretend otherwise. It's useful for the other kind of work: scheduled, unattended jobs where nobody is waiting on the cursor. My use is day/week/month end review — go through the numbers, flag what doesn't reconcile, draft the report — and there the two things that matter are that the model is good enough to trust with the judgement (K3 is, and it's the full 2.8T model, not a cut-down one) and that the data never leaves the machine.</p>
]]></description><pubDate>Tue, 08 Sep 2026 20:25:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49616520</link><dc:creator>Argonautlabs</dc:creator><comments>https://news.ycombinator.com/item?id=49616520</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49616520</guid></item></channel></rss>