<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: skohan</title><link>https://news.ycombinator.com/user?id=skohan</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 22 Sep 2026 01:12:23 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=skohan" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by skohan in "M5 Ultra Mac Studio Review"]]></title><description><![CDATA[
<p>1.2 T/s is not that modest is it?  That's very close to an RTX pro 5000</p>
]]></description><pubDate>Mon, 21 Sep 2026 20:46:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49793148</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49793148</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49793148</guid></item><item><title><![CDATA[New comment by skohan in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>I don't think smart people generally solve problems by talking through reasoning steps at a mile a minute.  They clear their mind and let the solution come.<p>Of course I don't know if there's really a way for this to be molded in current LLM's (sounds more like diffusion)</p>
]]></description><pubDate>Tue, 01 Sep 2026 20:12:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49527469</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49527469</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49527469</guid></item><item><title><![CDATA[New comment by skohan in "Apple caught off guard by AI demand for Mac Mini and Mac Studio"]]></title><description><![CDATA[
<p>16GB VRAM is probably not quite enough.  With 32GB, or <i>maybe</i> even 24GB, you can do serious coding work with local models.</p>
]]></description><pubDate>Tue, 01 Sep 2026 07:44:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49519187</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49519187</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49519187</guid></item><item><title><![CDATA[New comment by skohan in "Apple caught off guard by AI demand for Mac Mini and Mac Studio"]]></title><description><![CDATA[
<p>For decode, memory bandwidth is the main bottleneck, so these machines will likely perform well even without a ton of GPU horsepower.  Not as well as Blackwell, but I expect they will be a reasonable choice in terms of price/performance if you want to run large models with a lot of context.<p>The main place they are a bit behind is in the number formats they support natively.  Iirc M5 doesn't have native FP8 support, so you will take a speed penalty on quants where other architectures get better acceleration.</p>
]]></description><pubDate>Tue, 01 Sep 2026 07:42:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49519175</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49519175</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49519175</guid></item><item><title><![CDATA[New comment by skohan in "Asahi Linux Progress Report: Linux 7.2"]]></title><description><![CDATA[
<p>That's what I mean though.  The comment I was responding to was talking about "forever laptops" - my point is there's plenty of room for new capabilities which will make current hardware obsolete.  Just like how GPU's didn't exist at all, and became a standard part of computing.<p>And given how fast the hardware and software is evolving, I can easily imagine a future where we all have very capable models running on our own devices for an embedded intelligence layer that's doing most of the day-to-day tasks, and only have to outsource to a super-smart cloud model for specific things.</p>
]]></description><pubDate>Thu, 27 Aug 2026 17:25:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49468211</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49468211</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49468211</guid></item><item><title><![CDATA[New comment by skohan in "Asahi Linux Progress Report: Linux 7.2"]]></title><description><![CDATA[
<p>As someone who does a lot of work with local LLM's, today's systems feel woefully under-powered.  I'm looking forward to a future where my laptop has 10x the memory, 100x the memory bandwidth, and optimized cores to make inference workflows that currently take minutes or hours go down to seconds or milliseconds.<p>While we're at the point where traditional software is pretty much fast enough for all but extreme use-cases, with LLM's it feels like we're back to the days where you press compile and go have a coffee or chat to your colleague.</p>
]]></description><pubDate>Thu, 27 Aug 2026 13:05:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49464274</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49464274</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49464274</guid></item><item><title><![CDATA[New comment by skohan in "The Harness Is the Thing"]]></title><description><![CDATA[
<p>Ok that's fair if it's targeting a non-technical audience.<p>But I think this will eventually be a problem solved at the OS level in a more streamlined way.  I.e. there will be fine-grained permissions you need to approve to give an agent access to the system.</p>
]]></description><pubDate>Thu, 27 Aug 2026 12:57:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49464142</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49464142</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49464142</guid></item><item><title><![CDATA[New comment by skohan in "The Harness Is the Thing"]]></title><description><![CDATA[
<p>That's the most basic version of a harness.  But a harness is really about automating context delivery to the LLM based on your use-case.</p>
]]></description><pubDate>Thu, 27 Aug 2026 12:54:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49464097</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49464097</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49464097</guid></item><item><title><![CDATA[New comment by skohan in "The Harness Is the Thing"]]></title><description><![CDATA[
<p>Is that the harness' job?  It seems to me the best place for sandboxing is at the OS level (i.e. running the harness inside a container with correct access configured).</p>
]]></description><pubDate>Thu, 27 Aug 2026 12:41:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49463902</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49463902</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49463902</guid></item><item><title><![CDATA[New comment by skohan in "The Hugging Face incident and the road ahead"]]></title><description><![CDATA[
<p>And despite the enormous capital expenditure, Chinese models are nipping at their heels at what must be a fraction of the cost.  Sometimes constraints are healthy for inducing creative solutions.</p>
]]></description><pubDate>Thu, 27 Aug 2026 12:36:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49463817</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49463817</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49463817</guid></item><item><title><![CDATA[New comment by skohan in "Nvidia agrees to acquire Hugging Face for $13B"]]></title><description><![CDATA[
<p>Is huggingface profitable?<p>I'm not a fan of big-tech acquisition results either, but one benefit can be that a product continues to exist when it would otherwise become insolvent.</p>
]]></description><pubDate>Thu, 27 Aug 2026 06:40:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49460728</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49460728</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49460728</guid></item><item><title><![CDATA[New comment by skohan in "Nvidia agrees to acquire Hugging Face for $13B"]]></title><description><![CDATA[
<p>That could be a factor, but the optimistic interpretation would be that they want to support the open ecosystem because it sells more chips.<p>Models are <i>already</i> largely hardware agnostic.  It would be pretty hard to put that cat back in the bag.<p>I could imagine them building value-added services on top of HF to advantage Nvidia products (i.e. "run this model on NVIDIA cloud" with one-click), but in this moment it's hard to imagine how they could actively disadvantage models built to run on other platforms.</p>
]]></description><pubDate>Thu, 27 Aug 2026 06:37:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49460708</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49460708</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49460708</guid></item><item><title><![CDATA[New comment by skohan in "Nvidia agrees to acquire Hugging Face for $13B"]]></title><description><![CDATA[
<p>They didn't, but their price increases haven't kept pace with general memory price increases (yet) so currently their pricing seems reasonable.</p>
]]></description><pubDate>Thu, 27 Aug 2026 06:31:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49460657</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49460657</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49460657</guid></item><item><title><![CDATA[New comment by skohan in "Nvidia agrees to acquire Hugging Face for $13B"]]></title><description><![CDATA[
<p>Yeah I think this is probably one of the best outcomes you could hope for if you want to use open models.</p>
]]></description><pubDate>Thu, 27 Aug 2026 06:29:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49460642</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49460642</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49460642</guid></item><item><title><![CDATA[New comment by skohan in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>I'm using unsloth dynamic Q4_K_XL.<p>My use-case is coding, currently working on a project with a Rust backend and TS/React/Vite frontend, with probably tens of thousands of lines of code total (including tests).</p>
]]></description><pubDate>Thu, 27 Aug 2026 06:17:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49460546</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49460546</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49460546</guid></item><item><title><![CDATA[New comment by skohan in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>I'm using the unsloth dynamic Q4 and getting good results.  I was running Q5, but Q4 gives more context headroom so I can run two agents in parallel with ~100k context each with 32GB vram.</p>
]]></description><pubDate>Thu, 27 Aug 2026 06:09:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49460485</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49460485</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49460485</guid></item><item><title><![CDATA[New comment by skohan in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>3.8 doesn't have a minimal thinking mode, only low, medium and xhigh.</p>
]]></description><pubDate>Thu, 27 Aug 2026 03:42:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49459393</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49459393</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49459393</guid></item><item><title><![CDATA[New comment by skohan in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>That sounds like something is off - I'm using UD-Q4_K_XL on pi with xhigh thinking, and unless I'm vastly underestimating the complexity of the script that's the kind of task I would expect to take a couple of minutes (getting ~30t/s decode).  What server are you running, and are you using the recommended parameters from qwen/unsloth?</p>
]]></description><pubDate>Thu, 27 Aug 2026 03:39:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49459364</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49459364</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49459364</guid></item><item><title><![CDATA[New comment by skohan in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>Your assumption is incorrect.</p>
]]></description><pubDate>Wed, 26 Aug 2026 19:16:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49454327</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49454327</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49454327</guid></item><item><title><![CDATA[New comment by skohan in "GLM-5.3-Flash"]]></title><description><![CDATA[
<p>I'm using pi inside a self-made harness.  I've found going super lightweight with context (AGENTS.md is maybe 20 lines) and letting the model discover what it needs to gives the best results.</p>
]]></description><pubDate>Wed, 26 Aug 2026 19:14:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49454306</link><dc:creator>skohan</dc:creator><comments>https://news.ycombinator.com/item?id=49454306</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49454306</guid></item></channel></rss>