<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: djsjajah</title><link>https://news.ycombinator.com/user?id=djsjajah</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 06 Aug 2026 07:18:52 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=djsjajah" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by djsjajah in "Something is changing in the unit economics of software"]]></title><description><![CDATA[
<p>You have to power it all the time, but the amount of power it uses while it’s on will change by up to a few orders of magnitude depending on the gpu. It’s not uncommon for a gpu to be pulling just a couple of watts at idles and several hundred at full tilt.<p>So the only way it’s a fixed cost is if you don’t pay for power. If you only consider the cost of the power, it might still be cheaper paying for an api.</p>
]]></description><pubDate>Thu, 06 Aug 2026 03:09:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49191938</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=49191938</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49191938</guid></item><item><title><![CDATA[New comment by djsjajah in "Our position on open-weights models"]]></title><description><![CDATA[
<p>but what is the point?
A ban is supposed to make a certain thing less likely to occur. Does a ban of open source models do that? Presumably, the behavior you are trying to limit is the miss-use of these models but I don't know how many state sponsored hacking groups are going to give a ban a second thought.</p>
]]></description><pubDate>Tue, 28 Jul 2026 01:00:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49077905</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=49077905</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49077905</guid></item><item><title><![CDATA[New comment by djsjajah in "GPT-5.6 Sol Ultra will be in Codex"]]></title><description><![CDATA[
<p>I think they were making a joke. In the future, you might consider the advice you are giving as well as giving it.</p>
]]></description><pubDate>Tue, 07 Jul 2026 01:01:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=48812454</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48812454</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48812454</guid></item><item><title><![CDATA[New comment by djsjajah in "Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5"]]></title><description><![CDATA[
<p>Yes. Wait a day</p>
]]></description><pubDate>Wed, 01 Jul 2026 07:12:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=48743258</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48743258</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48743258</guid></item><item><title><![CDATA[New comment by djsjajah in "Good results fine tuning a local LLM like Qwen 3:0.6B to categorize questions"]]></title><description><![CDATA[
<p>Not with 800 examples. If you are going to consider an ngram model, I think you are better off getting a frontier llm to write you an absurd regex.</p>
]]></description><pubDate>Mon, 22 Jun 2026 02:46:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48625024</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48625024</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48625024</guid></item><item><title><![CDATA[New comment by djsjajah in "Apple boss Tim Cook says prices to rise due to memory chip costs"]]></title><description><![CDATA[
<p>Except, that won’t help. By the time a new fab is up and running, we will probably have a massive surplus.</p>
]]></description><pubDate>Thu, 18 Jun 2026 07:36:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48582070</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48582070</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48582070</guid></item><item><title><![CDATA[New comment by djsjajah in "Why Is Claude Turning into an a**Hole?"]]></title><description><![CDATA[
<p>You need to think this thought through all the way to the end. What it has said also influences what it will say. If it has consistently made combative responses, then the most likely thing to do is to continue to be combative.<p>I don't think there is any way back after the conversation takes a turn like that so there is no point in arguing with it. The only thing you can do is to fork the conversation before it made the first mistake and give it more context or tell it to look things up.</p>
]]></description><pubDate>Sun, 14 Jun 2026 23:52:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48534458</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48534458</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48534458</guid></item><item><title><![CDATA[New comment by djsjajah in "Can LLMs Beat Classical Hyperparameter Optimization Algorithms?"]]></title><description><![CDATA[
<p>It’s amusing that a lot of the agents have worked out that sampling doesn’t change ppl.</p>
]]></description><pubDate>Tue, 09 Jun 2026 22:17:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48468560</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48468560</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48468560</guid></item><item><title><![CDATA[New comment by djsjajah in "Apple reveals new AI architecture built around Google Gemini models"]]></title><description><![CDATA[
<p>I think what they mean by “now” is the stuff announced today.</p>
]]></description><pubDate>Mon, 08 Jun 2026 21:03:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48452026</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=48452026</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48452026</guid></item><item><title><![CDATA[New comment by djsjajah in "The AI revolution in math has arrived"]]></title><description><![CDATA[
<p>I don't follow. Can you explain how your comment is relevant to mine? It might help if you also explain how you interpreted my comment.</p>
]]></description><pubDate>Tue, 14 Apr 2026 22:59:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=47772549</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47772549</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47772549</guid></item><item><title><![CDATA[New comment by djsjajah in "The AI revolution in math has arrived"]]></title><description><![CDATA[
<p>You just failed the Turing test.</p>
]]></description><pubDate>Tue, 14 Apr 2026 04:44:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=47761344</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47761344</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47761344</guid></item><item><title><![CDATA[New comment by djsjajah in "Taking on CUDA with ROCm: 'One Step After Another'"]]></title><description><![CDATA[
<p>I have 2 of them. I would advise against if you want to run things like vllm. I have had the cards for months and I still have not been able to create a uv env with trl and vllm. For vllm, it’s works fine in docker for some models. With one gpu, gpt-oss 20b decoding at a cumulative 600-800tps with 32 concurrent requests depending on context length but I was getting trash performance out of qwen3.5 and Gemma4<p>If I were to do it again, I’d probably just get a dgx spark. I don’t think it’s been worth the hassle.</p>
]]></description><pubDate>Mon, 13 Apr 2026 11:26:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=47750516</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47750516</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47750516</guid></item><item><title><![CDATA[New comment by djsjajah in "Taking on CUDA with ROCm: 'One Step After Another'"]]></title><description><![CDATA[
<p>> or by the community<p>Hmmm</p>
]]></description><pubDate>Mon, 13 Apr 2026 07:11:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=47748722</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47748722</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47748722</guid></item><item><title><![CDATA[New comment by djsjajah in "Quantization from the Ground Up"]]></title><description><![CDATA[
<p>yes, but the difference between one model and one 4x larger is usually a lot more than that.<p>It is not a question of do a run Qwen 8b at bf16 or a quantized version. It more of a question of do I run Qwen 8b at full precision or do I run a quantized version of Qwen 27b.<p>You will find that you are usually better off with the larger model.</p>
]]></description><pubDate>Thu, 26 Mar 2026 01:35:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=47525710</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47525710</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47525710</guid></item><item><title><![CDATA[New comment by djsjajah in "Tinybox – Offline AI device 120B parameters"]]></title><description><![CDATA[
<p>trl.
give me a uv command to get that working.<p>But even in the amd stack things (like ck and aiter) consumer cards are not even second class citizens. They are a distance third at best.
If you just want to run vllm with the latest model, if you can get it running at all there are going to be paper cuts all along the way and even then the performance won't be close to what you could be getting out of the hardware.</p>
]]></description><pubDate>Sun, 22 Mar 2026 00:19:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=47473019</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47473019</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47473019</guid></item><item><title><![CDATA[New comment by djsjajah in "Attention Residuals"]]></title><description><![CDATA[
<p>No. It seems to me that the comment is objectively incorrect.
The original comment was talking about inference and from what I can tell, it is strictly going to run slower than the model trained to the same loss without this approach (it has "minimal overhead"). The main point is that you wont need to train that model for as long.</p>
]]></description><pubDate>Fri, 20 Mar 2026 21:34:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=47460910</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=47460910</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47460910</guid></item><item><title><![CDATA[New comment by djsjajah in "GLM-5: Targeting complex systems engineering and long-horizon agentic tasks"]]></title><description><![CDATA[
<p>That’s kind of a moot point. Even if none of those overheads existed you would still be getting a a fractions of the mfu. Models are fundamental limited by memory bandwidth even with best case scenarios of sft or prefill.<p>And what are you doing that I/O is a bottleneck?</p>
]]></description><pubDate>Thu, 12 Feb 2026 01:01:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=46983567</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=46983567</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46983567</guid></item><item><title><![CDATA[New comment by djsjajah in "Nvidia Stock Crash Prediction"]]></title><description><![CDATA[
<p>> including all previous experiments<p>How far back do you go? What about experiments into architecture features that didn’t make the cut? What about pre-transformer attention?<p>But more generally, why are you so sure that they team that built Gemini didn’t exclusively use TPUs while they were developing it?<p>I think that one of the reasons that Gemini caught up so quickly is because they have so much compute at fraction of the price of everyone else.</p>
]]></description><pubDate>Tue, 20 Jan 2026 21:23:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=46697922</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=46697922</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46697922</guid></item><item><title><![CDATA[New comment by djsjajah in "Command-line Tools can be 235x Faster than your Hadoop Cluster (2014)"]]></title><description><![CDATA[
<p>Not only can it be streamed, but lz4 will probably make things quicker.</p>
]]></description><pubDate>Sun, 18 Jan 2026 22:31:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=46672841</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=46672841</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46672841</guid></item><item><title><![CDATA[New comment by djsjajah in "Databases in 2025: A Year in Review"]]></title><description><![CDATA[
<p>You just ruined my day. The post makes it sound like gel is now dead. The post by Vercel does not give me much hope either [1]. Last commit on the gel repo was two weeks ago.<p>[1] <a href="https://vercel.com/blog/investing-in-the-python-ecosystem" rel="nofollow">https://vercel.com/blog/investing-in-the-python-ecosystem</a></p>
]]></description><pubDate>Mon, 05 Jan 2026 10:45:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=46497316</link><dc:creator>djsjajah</dc:creator><comments>https://news.ycombinator.com/item?id=46497316</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46497316</guid></item></channel></rss>