<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: hagen8</title><link>https://news.ycombinator.com/user?id=hagen8</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 17 Sep 2026 13:27:05 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=hagen8" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by hagen8 in "2026 Eclipse Webcams"]]></title><description><![CDATA[
<p>Mountains close to leon?</p>
]]></description><pubDate>Wed, 12 Aug 2026 12:32:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49271392</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49271392</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49271392</guid></item><item><title><![CDATA[New comment by hagen8 in "DeepSeek V4 Flash 0731"]]></title><description><![CDATA[
<p>Cached input tokens are what drives most costs.</p>
]]></description><pubDate>Fri, 07 Aug 2026 19:21:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49215108</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49215108</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49215108</guid></item><item><title><![CDATA[New comment by hagen8 in "Building an Advanced Agentic Harness"]]></title><description><![CDATA[
<p>Wrong. They are commonly used by millions.</p>
]]></description><pubDate>Wed, 05 Aug 2026 15:13:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49184114</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49184114</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49184114</guid></item><item><title><![CDATA[New comment by hagen8 in "Building an Advanced Agentic Harness"]]></title><description><![CDATA[
<p>Check out academic papers about:<p>1. Hierarchical skills, workflow, skill learning
2. Meta Harness, self-learning harnesses
3. Trace/trajectory representation
4. Common agentic benchmarks<p>But first more basic things like
5. Blog posts form anthropic
6. How Claude Code/PI/ Hermes!! agent works
7. Agent sessions/ Forking/ Hooks</p>
]]></description><pubDate>Wed, 05 Aug 2026 15:12:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49184096</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49184096</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49184096</guid></item><item><title><![CDATA[New comment by hagen8 in "When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation"]]></title><description><![CDATA[
<p>Check out <a href="https://agents-last-exam.org/" rel="nofollow">https://agents-last-exam.org/</a> there is still room for improvements!</p>
]]></description><pubDate>Tue, 04 Aug 2026 17:20:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49171940</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49171940</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49171940</guid></item><item><title><![CDATA[New comment by hagen8 in "Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone"]]></title><description><![CDATA[
<p>There are certain physical limits. Calculations need to be done. Either less calculations are necessary for the intelligence, or u accept less intelligence. But there is a limit in what u can do with specific hardware.</p>
]]></description><pubDate>Tue, 04 Aug 2026 10:29:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49166624</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49166624</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49166624</guid></item><item><title><![CDATA[New comment by hagen8 in "Elevators"]]></title><description><![CDATA[
<p>Most importantly, after entering the elevator. First press the close button and then the floor. That way u, safe the time of pressing a button as the door is already closing.</p>
]]></description><pubDate>Fri, 31 Jul 2026 19:35:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49127637</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49127637</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49127637</guid></item><item><title><![CDATA[New comment by hagen8 in "AI companies are shredding rare books"]]></title><description><![CDATA[
<p>Where are the sources for that?</p>
]]></description><pubDate>Mon, 27 Jul 2026 14:11:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49069999</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49069999</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49069999</guid></item><item><title><![CDATA[New comment by hagen8 in "IRGC claims it destroyed Amazon's Bahrain data center"]]></title><description><![CDATA[
<p>This is the claude code frontend-skill.</p>
]]></description><pubDate>Fri, 24 Jul 2026 14:15:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49036007</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=49036007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49036007</guid></item><item><title><![CDATA[New comment by hagen8 in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>Just switch the model, its not that much effort tbh. And u can also get a cheaper model than 2.5 lite for the same intelligence</p>
]]></description><pubDate>Tue, 21 Jul 2026 16:09:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48994232</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=48994232</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48994232</guid></item><item><title><![CDATA[New comment by hagen8 in "Human mathematicians are being outcounterexampled"]]></title><description><![CDATA[
<p>This will soon happen with theoretical physics, computer science, and everything which can be verified cheaply. Then, we will have long running projects augmented by agents for 2-4 years while AI companies are collecting data of human workflows. After that we will see AI being able to do those projects by themselves. This will lead to super fast human progress and cheap products. The price of things will be bound by energy and natural resources. Interesting times are ahead of us</p>
]]></description><pubDate>Tue, 21 Jul 2026 12:12:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48991225</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=48991225</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48991225</guid></item><item><title><![CDATA[New comment by hagen8 in "Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU"]]></title><description><![CDATA[
<p>Inference costs will go down massively once they use the upcoming GPUs. I estimated that a model like GLM5.2 will be around 0.03USD/M output tokens in 2 years when the Feynman GPUs will be available in 2028. And this did not even consider architectural efficiency improvements. In mid 2027 we will already see a 10x reduction once everyone has switched to the Ruby architecture.<p>It will be feasible for everyone to have 20 different agents running at all times. A new world is coming</p>
]]></description><pubDate>Wed, 15 Jul 2026 19:51:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48926198</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=48926198</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48926198</guid></item><item><title><![CDATA[New comment by hagen8 in "Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU"]]></title><description><![CDATA[
<p>Some ppl don't like to hear it. But I would assume that token costs when using an inference provider are cheaper than electricity of using locally.<p>If we just take into account output token generation for simplicity. With 5tps u get 18k tokens an hour. That would costs around 0.005USD from an inference provider.<p>I estimate that the server consumes probably around 500W during inference.<p>In Germany where 1kwh cost around 0.3USD, 18k tokens inferred locally would therefore cost 0.15USD which is 30x the costs of using an inference provider.<p>But for ppl who worry about their data, running locally might still be good. However, they should be aware, that it is much less efficient than using an inference provider.<p>The efficiency gap will also significantly increase as new GPUs will make inference much more efficient.<p>EDIT: I first thought it'd be 180k token, but thanks to someone mentioning in the comments, it is 18k. I guess with that, it will be tough unless u got electricity almost for free. Also, the inference providers are probably still using H200/H100 for those small models. Once they use GB300 or next year the new Ruby GPUs, inference will be cheaper by a factor of 30. By then, running local models will mostly be about privacy.</p>
]]></description><pubDate>Wed, 15 Jul 2026 18:03:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48924791</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=48924791</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48924791</guid></item><item><title><![CDATA[New comment by hagen8 in "GPT-5.6"]]></title><description><![CDATA[
<p>In my opinion Opus is waaayy better in agentic orchestration. It feels like it can natively deal with multiple subagents whereas gpt needs to be taught extensively.</p>
]]></description><pubDate>Fri, 10 Jul 2026 07:34:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48856870</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=48856870</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48856870</guid></item><item><title><![CDATA[New comment by hagen8 in "Herdr: Agent multiplexer that lives in your terminal"]]></title><description><![CDATA[
<p>This is way to complex... Why don't just use some harness which manages all that and give u a good UI?</p>
]]></description><pubDate>Mon, 29 Jun 2026 09:12:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48716754</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=48716754</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48716754</guid></item><item><title><![CDATA[New comment by hagen8 in "1M context is now generally available for Opus 4.6 and Sonnet 4.6"]]></title><description><![CDATA[
<p>Well, the question is what is contributing to the usage. Because as the context grows, the amount of input tokens are increasing. A model call with 800K token as input is 8 times more expensive than a model call with 100K tokens as input. Especially if we resume a conversation and caching does not hit, it would be very expensive with API pricing.</p>
]]></description><pubDate>Sat, 14 Mar 2026 01:07:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=47372204</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=47372204</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47372204</guid></item><item><title><![CDATA[New comment by hagen8 in "1M context is now generally available for Opus 4.6 and Sonnet 4.6"]]></title><description><![CDATA[
<p>Did u use the API or subscription?</p>
]]></description><pubDate>Sat, 14 Mar 2026 00:57:57 +0000</pubDate><link>https://news.ycombinator.com/item?id=47372134</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=47372134</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47372134</guid></item><item><title><![CDATA[New comment by hagen8 in "Redox OS has adopted a Certificate of Origin policy and a strict no-LLM policy"]]></title><description><![CDATA[
<p>They will sooner or later change that policy or get very slow in keeping up.</p>
]]></description><pubDate>Tue, 10 Mar 2026 10:03:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=47321197</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=47321197</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47321197</guid></item><item><title><![CDATA[New comment by hagen8 in "GPT-5.4"]]></title><description><![CDATA[
<p>But does it use the same agent harness? Because the harness determines the behavior a lot.</p>
]]></description><pubDate>Fri, 06 Mar 2026 06:42:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=47271707</link><dc:creator>hagen8</dc:creator><comments>https://news.ycombinator.com/item?id=47271707</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47271707</guid></item></channel></rss>