<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: pimeys</title><link>https://news.ycombinator.com/user?id=pimeys</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 04 Aug 2026 02:52:12 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=pimeys" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by pimeys in "Qwen3.8-Max: A New Bar for Coding and Cowork"]]></title><description><![CDATA[
<p>They are profitable and active on the enterprise local model territory. You can RL a model with them for your own purposes and I heard good things about it.</p>
]]></description><pubDate>Mon, 03 Aug 2026 06:28:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49151934</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49151934</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49151934</guid></item><item><title><![CDATA[New comment by pimeys in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>Yeah, been running too many evals in the past week I start to mix the versions up. Probably should sleep...<p>We run an agent company and outside coding the new Gemini 3.6 Flash and GPT 5.6 Luna are very interesting. Luna can do a bit of research and create reports. Gemini is great for computer use.<p>For programming it's all Kimi K3 now.</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:39:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115451</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49115451</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115451</guid></item><item><title><![CDATA[New comment by pimeys in "Advancing the price-performance frontier with GPT‑5.6"]]></title><description><![CDATA[
<p>If you have agents and users, you can run evals and see how far the models go. Luna is not greatest in tool calls, but if you define your problem well and the tools well, it is comparable to Gemini 4 Flash with much lower price tag.</p>
]]></description><pubDate>Thu, 30 Jul 2026 19:10:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49114318</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49114318</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49114318</guid></item><item><title><![CDATA[New comment by pimeys in "Hamburg's Stadtpark: A Park Built to Be Used"]]></title><description><![CDATA[
<p>I don't remember can you swim in Lietzensee but you definitely can in Weissensee and I consider both to be quite central still.<p>But you are correct, there's not enough swimmable lakes in Berlin and when the weather is nice you better drive somewhere in Brandenburg if you don't want to be stuck in a fully packed pond with 1000 others.</p>
]]></description><pubDate>Thu, 30 Jul 2026 11:17:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49108443</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49108443</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49108443</guid></item><item><title><![CDATA[New comment by pimeys in "Using an open model feels surprisingly good"]]></title><description><![CDATA[
<p>Yes. I tested OpenClaw that burned 15€ just by starting it and I hated its configuration. Spent 20€ in tokens and built my own in a day, that does everything I want and sips tokens.<p>What a time to be alive.</p>
]]></description><pubDate>Wed, 29 Jul 2026 06:43:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49094130</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49094130</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49094130</guid></item><item><title><![CDATA[New comment by pimeys in "Using an open model feels surprisingly good"]]></title><description><![CDATA[
<p>Funny. I use Fable a lot at work, but this one I paid from my own pocket and coded it with GLM 5.2.<p>Paid 20 euros in total.</p>
]]></description><pubDate>Tue, 28 Jul 2026 15:41:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49085574</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49085574</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49085574</guid></item><item><title><![CDATA[New comment by pimeys in "Using an open model feels surprisingly good"]]></title><description><![CDATA[
<p>I have these around the house:<p><a href="https://www.home-assistant.io/voice-pe/" rel="nofollow">https://www.home-assistant.io/voice-pe/</a><p>Then you can set the background AI to be any OpenAI compatible API. So I just created one to my local Rust Agent, and connected it to home assistant. Now I can yell from the couch to create me a new Proxmox container with the next free static IP address etc. :D</p>
]]></description><pubDate>Tue, 28 Jul 2026 15:40:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49085550</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49085550</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49085550</guid></item><item><title><![CDATA[New comment by pimeys in "Using an open model feels surprisingly good"]]></title><description><![CDATA[
<p>I wrote my own OpenClaw one weekend and I am running it as my assistant through Matrix with DeepSeek v4 Flash (and Qwen). It probably costs me about 2 dollars a month and is even more useful than ChatGPT would be due to me having full control on what tools it has access to.<p>I can do things like take a photo of a doctor's note among add the appointment to my calendar, send a PDF to my archive tagged, OCR'd etc, search info from internet, look data from Google Maps.<p>I can even integrate this to home assistant and talk to my agent with my open source Alexa-like system.</p>
]]></description><pubDate>Tue, 28 Jul 2026 03:34:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49078988</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49078988</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49078988</guid></item><item><title><![CDATA[New comment by pimeys in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>I've used GLM-5.2 a lot on fireworks and had never ever issues on rate limits. If they cannot handle the load with K3, there's the priority tier to get your evals done.<p>I'm definitely having full eval suite on already if they get overloaded later on.</p>
]]></description><pubDate>Mon, 27 Jul 2026 16:31:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49071997</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49071997</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49071997</guid></item><item><title><![CDATA[New comment by pimeys in "Kimi-K3 on HuggingFace"]]></title><description><![CDATA[
<p>Depends what you do. We have certain tasks we spend money on where Gemini 4.6 definitely is better than Opus 5.</p>
]]></description><pubDate>Mon, 27 Jul 2026 14:31:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49070270</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49070270</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49070270</guid></item><item><title><![CDATA[New comment by pimeys in "Elevated errors on Claude Opus 5"]]></title><description><![CDATA[
<p>I remember when 4.7 and 4.8 were released and people were asking what's wrong with them and 4.6 is the best.<p>But yes, I also think it's not the greatest model for programming. On the other hand, for agentic tasks that are not programming related it's hard to beat Opus 4.8. It can try different things and pivot even when the user is not great with prompting. 5.0 seems to not be worse, but definitely wastes more tokens and costs more.</p>
]]></description><pubDate>Mon, 27 Jul 2026 13:12:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49069236</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=49069236</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49069236</guid></item><item><title><![CDATA[New comment by pimeys in "Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling"]]></title><description><![CDATA[
<p>Depending on the harness, $200/session can be common. Thanks for harnesses like maki the cost is lower compared to OpenCode or Cursor.<p>But $200/day is very common in the business. The monthly plans are just a trial version of what's going to come for all of us. That's why I am very carefully looking into the open models today.</p>
]]></description><pubDate>Mon, 20 Jul 2026 21:28:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=48985109</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48985109</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48985109</guid></item><item><title><![CDATA[New comment by pimeys in "The lost joy of music piracy"]]></title><description><![CDATA[
<p>Ty</p>
]]></description><pubDate>Thu, 16 Jul 2026 21:52:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48940769</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48940769</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48940769</guid></item><item><title><![CDATA[New comment by pimeys in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>I use V4 flash as my personal agent. It categorizes documents, organizes my calendar, searches information etc. for pennies. Amazing model.<p>Not very good for programming though.</p>
]]></description><pubDate>Thu, 16 Jul 2026 21:31:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=48940528</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48940528</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48940528</guid></item><item><title><![CDATA[New comment by pimeys in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>GLM has issues with tool calls and nested JSON and it wastes tokens pretty often. I see it being a bit above half the price of Opus in a bit more complex eval tasks. With some RL you could probably get the tool calls sorted and the price down.</p>
]]></description><pubDate>Thu, 16 Jul 2026 21:28:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48940490</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48940490</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48940490</guid></item><item><title><![CDATA[New comment by pimeys in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>This is a question I am going to get an answer tomorrow with evals. Extremely interesting...</p>
]]></description><pubDate>Thu, 16 Jul 2026 21:26:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48940471</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48940471</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48940471</guid></item><item><title><![CDATA[New comment by pimeys in "WorkOS Having a Major Outage"]]></title><description><![CDATA[
<p>One hour and no updates. Still down.</p>
]]></description><pubDate>Thu, 16 Jul 2026 09:24:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=48932187</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48932187</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48932187</guid></item><item><title><![CDATA[New comment by pimeys in "GPT-5.6"]]></title><description><![CDATA[
<p>For me the weirdest is Luna. It costs the same as GLM to solve a task with it, but it just calls the same failing tool over and over again until we cut it out.<p>Now if you look at where GLM stands, or even DeepSeek v4 Flash, things get really interesting for what they provide.<p>US labs completely miss a cheap model that can solve problems for 95% of the people. Gemini 3.5 Flash could've been it if it didn't burn so many tokens.</p>
]]></description><pubDate>Fri, 10 Jul 2026 22:55:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48866361</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48866361</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48866361</guid></item><item><title><![CDATA[New comment by pimeys in "GPT-5.6"]]></title><description><![CDATA[
<p>I've been testing Sol/Terra/Luna now since yesterday, running complex evals on all of them and I feel a bit... mixed on how they perform.<p>The eval is an agent that runs a set of tools and a prompt we can tune separately for different models. The OpenAI version of the prompt was specifically tuned based on their guide[0]. Then we let Opus to run another agent that acts as a user, trying to solve a problem (anonymized and taken from production). The problem is complex and we don't expect it to be solved by these agents, but we measure how the agents operate when faced with a vague problem:<p>- Opus 4.8 and GLM 5.2 both identified a constraint sooner and stopped so the user can fix an issue first that the agent cannot solve.<p>- Sol tried hard to solve the issue with different tools, burning tokens, until finally reached to the same conclusion with Opus and GLM. It was two times more expensive compared to Opus and six times more expensive to GLM for this task.<p>- Terra went even further and started calling tools that would not solve the issue, burning tokens and failing.<p>- Luna repeated the same failing tool call until it hit the round limit, and burned more money than Opus.<p>I'm kind of puzzled with the new GPT. Like, yes Sol is OK for programming, but I was expecting to get a cheap agentic model for non-programming tasks, one that can detect if things go awry and correct. Terra is too expensive and Luna not really fit for the task. Sonnet 5 is a bit better but more expensive than Opus 4.8, which is still the best in my evals. GLM 5.2 is extremely good if you can define the task and the tools clearly for it, and costs pennies!<p>[0] <a href="https://developers.openai.com/api/docs/guides/latest-model" rel="nofollow">https://developers.openai.com/api/docs/guides/latest-model</a></p>
]]></description><pubDate>Fri, 10 Jul 2026 13:02:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48859325</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48859325</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48859325</guid></item><item><title><![CDATA[New comment by pimeys in "Introducing Muse Spark 1.1"]]></title><description><![CDATA[
<p>I use it all the time through Fireworks. The normal version when I pay it myself and the fast one when company pays. It's really fast and I never get rate limited with my daily use.</p>
]]></description><pubDate>Thu, 09 Jul 2026 15:47:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48847875</link><dc:creator>pimeys</dc:creator><comments>https://news.ycombinator.com/item?id=48847875</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48847875</guid></item></channel></rss>