<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: woctordho</title><link>https://news.ycombinator.com/user?id=woctordho</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 02 Oct 2026 20:08:40 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=woctordho" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by woctordho in "CUDA for AMD on Windows"]]></title><description><![CDATA[
<p>No, 10.1 is current</p>
]]></description><pubDate>Mon, 14 Sep 2026 06:10:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49692612</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49692612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49692612</guid></item><item><title><![CDATA[New comment by woctordho in "Detecting and countering misuse of AI: September 2026"]]></title><description><![CDATA[
<p>Let me put my two cents: In China we've got accustomed to the fact that every word we say will be seen by the surveillance, so it's not a big problem that Anthropic also see it. Also we know that they can see it but they can't stop it. There are all kinds of ways to work around account blocking.<p>As the old saying goes, communists disdain to conceal their views and aims.</p>
]]></description><pubDate>Fri, 11 Sep 2026 05:39:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49653951</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49653951</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49653951</guid></item><item><title><![CDATA[New comment by woctordho in "Qwen 3.8 follows GPT-5.5 Pro reasoning prefills"]]></title><description><![CDATA[
<p>Reasoning works as long as there is a consistent latent space representation. Any kind of poison will just become part of the representation. There's evidence that even directly training on encrypted reasoning traces works, because the length is already a strong signal.</p>
]]></description><pubDate>Wed, 09 Sep 2026 21:35:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49634741</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49634741</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49634741</guid></item><item><title><![CDATA[New comment by woctordho in "Path to Astra: critical capabilities and frontier safeguards"]]></title><description><![CDATA[
<p>There are 'transfer stations' and that's how exactly I use GPT and Claude in China. OpenAI and Anthropic do not sell in China, so we use their AI with a much lower price like 1% of the official API price. The largest transfer stations have TBs of traffic every day, and the traffic is eventually possessed by the open source community.<p>Subscription engineering is a deep field. Neither OpenAI nor Anthropic have any technical advantage in this field.</p>
]]></description><pubDate>Wed, 02 Sep 2026 23:44:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49544132</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49544132</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49544132</guid></item><item><title><![CDATA[New comment by woctordho in "Path to Astra: critical capabilities and frontier safeguards"]]></title><description><![CDATA[
<p>All the RL data are exactly public. There are huge amount of distilled data freely available, and that amount is more than enough to train a ~10T model.</p>
]]></description><pubDate>Wed, 02 Sep 2026 11:02:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49534549</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49534549</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49534549</guid></item><item><title><![CDATA[New comment by woctordho in "Seedance 2.5"]]></title><description><![CDATA[
<p>MiniMax H3 is going to release weights. You can locally run it with definitely less than $10k (and possibly faster than Seedance's queue), and it's fun to train it for whatever you need.</p>
]]></description><pubDate>Sun, 02 Aug 2026 00:58:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49140072</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49140072</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49140072</guid></item><item><title><![CDATA[New comment by woctordho in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>GGUF is at least better than bnb. From what I know, bnb does not yet find a way to quantize MoE with enough accuracy, and maintain the dequant-MoE kernels. In the age of Qwen 3.0, people tried to make some bnb '4-bit' quants of MoE models, but actually the MoE part is not quantized. It's a pity that even Unsloth gave up low-VRAM finetuning with MoE (although they're making their GGUFs for inference), and the world of local training looks stagnated for months.<p>GGUF is maintained by all the llama.cpp developers. There are many quantization formats and algorithms under this container format, some are optimized for MoE (such as APEX quant), some for CPU and some for GPU, some work surprisingly well below 4-bit (and even near 1-bit). It also supports recent architectures like linear attentions and mHC.</p>
]]></description><pubDate>Mon, 27 Jul 2026 14:59:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49070663</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49070663</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49070663</guid></item><item><title><![CDATA[New comment by woctordho in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.<p>I've made some proof of concept in <a href="https://github.com/woct0rdho/transformers5-qwen3.5-recipe" rel="nofollow">https://github.com/woct0rdho/transformers5-qwen3.5-recipe</a> . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.<p>Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.</p>
]]></description><pubDate>Mon, 27 Jul 2026 07:45:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49066303</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49066303</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49066303</guid></item><item><title><![CDATA[New comment by woctordho in "Petals: Run LLMs at home, BitTorrent-style"]]></title><description><![CDATA[
<p>Distributed training is much harder than distributed inference but not impossible. See the recent development of DiLoCo at Nous Research and Prime Intellect.</p>
]]></description><pubDate>Thu, 23 Jul 2026 04:23:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49016904</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49016904</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49016904</guid></item><item><title><![CDATA[New comment by woctordho in "Petals: Run LLMs at home, BitTorrent-style"]]></title><description><![CDATA[
<p>Relevant: Why Switzerland has 25 Gbit internet and America doesn't <a href="https://news.ycombinator.com/item?id=47652400">https://news.ycombinator.com/item?id=47652400</a></p>
]]></description><pubDate>Thu, 23 Jul 2026 02:36:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49016176</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49016176</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49016176</guid></item><item><title><![CDATA[New comment by woctordho in "Petals: Run LLMs at home, BitTorrent-style"]]></title><description><![CDATA[
<p>AI Horde has some measures to prevent Sybil attack that returns wrong results, but not enforce zero data retention. Prompts belong to the whole open source community. For example <a href="https://huggingface.co/datasets/la-ji/sd-prompt-in-the-wild" rel="nofollow">https://huggingface.co/datasets/la-ji/sd-prompt-in-the-wild</a></p>
]]></description><pubDate>Thu, 23 Jul 2026 02:34:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49016156</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49016156</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49016156</guid></item><item><title><![CDATA[New comment by woctordho in "Petals: Run LLMs at home, BitTorrent-style"]]></title><description><![CDATA[
<p>Petals is from 2022. Nowadays intelligence of smaller models, quantization techs, and optimizations to run models faster on  consumer GPUs have improved a lot.<p>For distributed inference of smaller LLMs and diffusion models that fits in one consumer GPU rather than splits on multiple machines, there are already pretty good solutions such as AI Horde (formerly Stable Horde) [0]. Notably, it's the default provider that powers SillyTavern. It also has an interesting economy model of kudos.<p>[0] <a href="https://stablehorde.net/" rel="nofollow">https://stablehorde.net/</a></p>
]]></description><pubDate>Thu, 23 Jul 2026 02:31:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49016123</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49016123</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49016123</guid></item><item><title><![CDATA[New comment by woctordho in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>So is making a PR different from making the whole software. This is what an open source community is good for.</p>
]]></description><pubDate>Wed, 22 Jul 2026 09:26:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49003939</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=49003939</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49003939</guid></item><item><title><![CDATA[New comment by woctordho in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>See the recent development of DiLoCo at Nous Research and Prime Intellect.</p>
]]></description><pubDate>Tue, 21 Jul 2026 11:17:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=48990751</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48990751</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48990751</guid></item><item><title><![CDATA[New comment by woctordho in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>There's a lot of individual effort of improving the models. See how many finetuned models and LoRAs are there on Hugging Face.</p>
]]></description><pubDate>Tue, 21 Jul 2026 11:15:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48990745</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48990745</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48990745</guid></item><item><title><![CDATA[New comment by woctordho in "Qwen 3.8"]]></title><description><![CDATA[
<p>There is a forum named Zhihu. AI translation works mostly well to translate contents there into English.</p>
]]></description><pubDate>Tue, 21 Jul 2026 10:55:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=48990609</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48990609</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48990609</guid></item><item><title><![CDATA[New comment by woctordho in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>Yes in a mid-sized company. I'm exactly doing this, and what I'm competing against is the OpenAI API priced 0.2 CNY = 1 USD in China.</p>
]]></description><pubDate>Fri, 17 Jul 2026 04:10:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=48943256</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48943256</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48943256</guid></item><item><title><![CDATA[New comment by woctordho in "Alternative(s) to run CUDA on non-Nvidia hardware"]]></title><description><![CDATA[
<p>There's nothing wrong to run CUDA on non-Nvidia hardware. CUDA has an interface that is reasonably well-designed, well-documented/reverse-engineered, and battle-tested for decades. What we need is not to invent another interface just under the name of 'open standard', but to implement the same interface. ROCm is exactly doing this, and so are other hardware SDKs such as MooreThread and Alibaba T-Head.</p>
]]></description><pubDate>Tue, 14 Jul 2026 09:41:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48904271</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48904271</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48904271</guid></item><item><title><![CDATA[New comment by woctordho in "DSpark: Speculative decoding accelerates LLM inference [pdf]"]]></title><description><![CDATA[
<p>And humans don't run on markets.</p>
]]></description><pubDate>Sat, 27 Jun 2026 12:06:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48697534</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48697534</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48697534</guid></item><item><title><![CDATA[New comment by woctordho in "The gap between open weights LLMs and closed source LLMs"]]></title><description><![CDATA[
<p>Fun fact: Hacker News is canonically banned in China, but I'm still talking here. There are plenty of techs to work around region block. The incentive to report somebody is comically called '50w' (500k CNY) and no one gives a shit about it in real life.</p>
]]></description><pubDate>Sat, 27 Jun 2026 06:54:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48695828</link><dc:creator>woctordho</dc:creator><comments>https://news.ycombinator.com/item?id=48695828</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48695828</guid></item></channel></rss>