<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: zackangelo</title><link>https://news.ycombinator.com/user?id=zackangelo</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 04 Sep 2026 09:54:47 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=zackangelo" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by zackangelo in "Qwen 3.8 27B available on Cerebras at 1500 tokens/s"]]></title><description><![CDATA[
<p>Something a lot of model providers don't talk about: any time an engine uses speculative decoding the throughput will depend on how much your output token distribution matches what the draft model was trained on.<p>The DFlash2 draft model we're using was trained on a lot of code, so if you use it in a coding agent you'll probably notice it run a lot faster (we've seen it break 300 tok/s).</p>
]]></description><pubDate>Thu, 03 Sep 2026 20:33:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49556501</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=49556501</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49556501</guid></item><item><title><![CDATA[New comment by zackangelo in "Qwen 3.8 27B available on Cerebras at 1500 tokens/s"]]></title><description><![CDATA[
<p>just added 8 more H200s to the cluster, if you (or anyone else) runs into issues please feel free to drop me a message: zack at mixlayer.com</p>
]]></description><pubDate>Thu, 03 Sep 2026 20:15:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49556217</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=49556217</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49556217</guid></item><item><title><![CDATA[New comment by zackangelo in "Qwen 3.8 27B available on Cerebras at 1500 tokens/s"]]></title><description><![CDATA[
<p>apologies we just got a sudden burst of new users and traffic, it's scaling up now.</p>
]]></description><pubDate>Thu, 03 Sep 2026 19:34:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49555507</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=49555507</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49555507</guid></item><item><title><![CDATA[New comment by zackangelo in "Qwen 3.8 27B available on Cerebras at 1500 tokens/s"]]></title><description><![CDATA[
<p>We're serving it around 150-200tok/s (uses our new speculative decoding implementation on a DFlash2 draft model).<p><a href="https://mixlayer.com" rel="nofollow">https://mixlayer.com</a>, LAUNCH-Q38-27B gets you $5 in credits if you want to kick the tires.</p>
]]></description><pubDate>Thu, 03 Sep 2026 19:05:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49555046</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=49555046</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49555046</guid></item><item><title><![CDATA[New comment by zackangelo in "Rivian OS 2"]]></title><description><![CDATA[
<p>What was wrong with your R1S?<p>I had an R1T Launch Edition for a little over a year and it was hands down one of the best cars I’ve ever owned. The only issue I had was with the fob proximity sensor. The only reason I sold it was it was a bit too big for me.</p>
]]></description><pubDate>Thu, 20 Aug 2026 16:35:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49376941</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=49376941</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49376941</guid></item><item><title><![CDATA[New comment by zackangelo in "DFlash 2: Keep Drafting Parallel"]]></title><description><![CDATA[
<p>DFlash is lossless so this would be a bug in the implementation if it is indeed a regression against the target model.</p>
]]></description><pubDate>Thu, 20 Aug 2026 00:31:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49369007</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=49369007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49369007</guid></item><item><title><![CDATA[New comment by zackangelo in "Performance per dollar is getting faster and cheaper"]]></title><description><![CDATA[
<p>Blackwell supports nvfp4 natively.</p>
]]></description><pubDate>Sat, 04 Jul 2026 02:25:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48782133</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48782133</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48782133</guid></item><item><title><![CDATA[New comment by zackangelo in "Two Qwen3 models on one DGX Spark: the residency math"]]></title><description><![CDATA[
<p>what was the concurrency limitation? that node should be able to support a lot more</p>
]]></description><pubDate>Sun, 21 Jun 2026 17:26:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=48620755</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48620755</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48620755</guid></item><item><title><![CDATA[New comment by zackangelo in "Kimi K2.7-Code: open-source coding model with better token efficiency"]]></title><description><![CDATA[
<p>I don't believe safetensors has a native int4 dtype, so they packed 4 int4s into a bf16 in this checkpoint.</p>
]]></description><pubDate>Fri, 12 Jun 2026 17:51:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48507268</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48507268</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48507268</guid></item><item><title><![CDATA[New comment by zackangelo in "The real cost of owning a home"]]></title><description><![CDATA[
<p>If you're in SF and weighing this decision, it's easy to get tilted in the buy direction because the rental stock is so horrific. Landlords have very little incentive to update properties or provide basic amenities that people take for granted in other major cities (good luck getting a washer/dryer).</p>
]]></description><pubDate>Tue, 26 May 2026 16:58:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48282437</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48282437</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48282437</guid></item><item><title><![CDATA[New comment by zackangelo in "Qwen3.7-Max: The Agent Frontier"]]></title><description><![CDATA[
<p>With the 3.5 release, the Plus model was just a rebrand of the open weight 397B. But I suspect that will change going forward. They haven’t released the weights for 3.6 but they did make it available through a few US providers.</p>
]]></description><pubDate>Wed, 20 May 2026 15:09:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48209099</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48209099</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48209099</guid></item><item><title><![CDATA[New comment by zackangelo in "I’ve joined Anthropic"]]></title><description><![CDATA[
<p>absolutely not, take Kimi K2.6 for a spin</p>
]]></description><pubDate>Tue, 19 May 2026 17:29:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48196381</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48196381</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48196381</guid></item><item><title><![CDATA[How do agents see your website?]]></title><description><![CDATA[
<p>Article URL: <a href="https://what-do-agents-see.runtype.app/">https://what-do-agents-see.runtype.app/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48123838">https://news.ycombinator.com/item?id=48123838</a></p>
<p>Points: 4</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 13 May 2026 16:11:19 +0000</pubDate><link>https://what-do-agents-see.runtype.app/</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=48123838</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48123838</guid></item><item><title><![CDATA[New comment by zackangelo in "Mistral Medium 3.5"]]></title><description><![CDATA[
<p>Isn't Kimi K2.6 natively INT4?</p>
]]></description><pubDate>Wed, 29 Apr 2026 17:56:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=47951935</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=47951935</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47951935</guid></item><item><title><![CDATA[New comment by zackangelo in "HashiCorp co-founder says GitHub 'no longer a place for serious work'"]]></title><description><![CDATA[
<p>I don’t think this is true across Blizzard. Overwatch is the best it’s ever been.</p>
]]></description><pubDate>Wed, 29 Apr 2026 12:54:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=47947665</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=47947665</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47947665</guid></item><item><title><![CDATA[New comment by zackangelo in "Parallel agents in Zed"]]></title><description><![CDATA[
<p>I give them a try about twice a year. I write a lot of Rust which should be squarely in their wheelhouse.<p>This last time I was pleasantly surprised to find they mostly fixed their SSH remote editing support. But then it started truncating rustc inline error messages and I couldn’t figure out how to view the whole thing easily. When you’re just trying to get something done little bits like this can add up quickly. Punted back to Cursor for now.</p>
]]></description><pubDate>Wed, 22 Apr 2026 23:35:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=47870612</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=47870612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47870612</guid></item><item><title><![CDATA[New comment by zackangelo in "Qwen3.6-35B-A3B: Agentic coding power, now open to all"]]></title><description><![CDATA[
<p>They are but the IDE needs to be integrated with them.<p>Qwen specifically calls out FIM (“fill in the middle”) support on the model card and you can see it getting confused and posting the control tokens in the example here.</p>
]]></description><pubDate>Thu, 16 Apr 2026 15:00:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=47794134</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=47794134</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47794134</guid></item><item><title><![CDATA[New comment by zackangelo in "Qwen3.6-35B-A3B: Agentic coding power, now open to all"]]></title><description><![CDATA[
<p>17b per token. So when you’re generating a single stream of text (“decoding”) 17b parameters are active.<p>If you’re decoding multiple streams, it will be 17b per stream (some tokens will use the same expert, so there is some overlap).<p>When the model is ingesting the prompt (“prefilling”) it’s looking at many tokens at once, so the number of active parameters will be larger.</p>
]]></description><pubDate>Thu, 16 Apr 2026 14:57:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=47794079</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=47794079</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47794079</guid></item><item><title><![CDATA[New comment by zackangelo in "GPU memory snapshots: sub-second startup (2025)"]]></title><description><![CDATA[
<p>This uses Nvidia’s CUDA snapshot API under the hood, but you have to pair it with a host side snapshot as well. Modal uses gVisor for this, which is notoriously high overhead.<p>Does anyone know of a more efficient alternative if you’re running a trusted container?</p>
]]></description><pubDate>Sat, 10 Jan 2026 23:56:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=46571234</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=46571234</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46571234</guid></item><item><title><![CDATA[New comment by zackangelo in "macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt"]]></title><description><![CDATA[
<p>You’re right I misunderstood.<p>I’m not sure if it would be of much utility because this would presumably be for tensor parallel workloads. In that case you want the ranks in your cluster to be uniform or else everything will be forced to run at the speed of the slowest rank.<p>You could run pipeline parallel but not sure it’d be that much better than what we already have.</p>
]]></description><pubDate>Fri, 12 Dec 2025 23:26:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=46250310</link><dc:creator>zackangelo</dc:creator><comments>https://news.ycombinator.com/item?id=46250310</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46250310</guid></item></channel></rss>