<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: easygenes</title><link>https://news.ycombinator.com/user?id=easygenes</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Mon, 27 Jul 2026 21:36:07 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=easygenes" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by easygenes in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>That’s not what this indicates. This is the biggest and most expensive to serve, and most capable open weights model yet. They’re just pricing it in line with capabilities.<p>Kimi also offers generous subscriptions. Subs aren’t going anywhere. Think of subs like running an insurance business. There might be some users you lose money on (ones who max out their weekly quota without fail), but they’re managed such that the average subscription turns a healthy profit. There’s never been subsidies in model serving, inference is just cheaper in terms of ops TCO than people assume, and API margins are very high.</p>
]]></description><pubDate>Thu, 16 Jul 2026 16:00:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=48936338</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48936338</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48936338</guid></item><item><title><![CDATA[New comment by easygenes in "Leanstral 1.5: Proof abundance for all"]]></title><description><![CDATA[
<p>Was fun to see their developers make nods to Le Chaton Fat in the announcements for this on Twitter.<p>I suspect a true "big new general-purpose" model is around the corner from them, whether or not they were in on Le Chaton Fat for real. They've mentioned it after the media circus. Hopefully more creatively named than just "Large 4".</p>
]]></description><pubDate>Sat, 04 Jul 2026 07:01:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48783240</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48783240</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48783240</guid></item><item><title><![CDATA[New comment by easygenes in "Claude Sonnet 5"]]></title><description><![CDATA[
<p>I'm a heavy enough user that I have both the OAI and Anth $200 plans. I always use at least 50% of my weekly Opus quota at Extra setting (meaning I use double the limit of the $100 plan, at minimum). Max I rarely touch because it is twice as slow and the incremental capability gain is minimal. Usually if Opus can't sort something well at Extra, the answer isn't to use Max but to hand the issue off to GPT-5.5 at XHigh.</p>
]]></description><pubDate>Tue, 30 Jun 2026 23:37:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48740650</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48740650</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48740650</guid></item><item><title><![CDATA[New comment by easygenes in "HackerRank open sourced its ATS. My resume scored 90/100. Oh wait 74. No – 88"]]></title><description><![CDATA[
<p>There are. If the kernels are nondeterministic (e.g. timing issues) there are minor changes between runs, on a single system, even with eager decode enabled (typically what temperature=0 achieves).</p>
]]></description><pubDate>Mon, 29 Jun 2026 06:12:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48715458</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48715458</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48715458</guid></item><item><title><![CDATA[New comment by easygenes in "Previewing GPT‑5.6 Sol: a next-generation model"]]></title><description><![CDATA[
<p>This is a strange one. We know the hardware capabilities of Cerebras force them to do aggressive REAP pruning to serve Kimi K2.6. Meaning that about 750B parameters is the upper limit of what they can serve economically. Not sure if this means Sol is smaller than anyone thinks or that they're just going to charge so much that a very inefficient serving regime is feasible.</p>
]]></description><pubDate>Sat, 27 Jun 2026 06:13:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48695652</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48695652</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48695652</guid></item><item><title><![CDATA[New comment by easygenes in "GLM-5.2 – How to Run Locally"]]></title><description><![CDATA[
<p>M5 Ultra will ship before end of year, likely. Though with current RAM shortage, likely max spec will be 256GB and in short supply.<p>In late 2027 or early 2028, Nvidia will release Vera Rubin DGX Spark, likely with double or better the performance of current Blackwell, though unclear if memory capacity will go up much from current 128GB. Two to four of those will run models like this decently.<p>In 2028 we should expect Vera Rubin RTX discrete lineup, including the replacement to the RTX PRO 6000. Likely memory spec will be minimum 128GB. Good chance of up to 200GB. Two to four of those will run NVFP4 models in this class very well.</p>
]]></description><pubDate>Tue, 23 Jun 2026 00:49:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48638673</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48638673</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48638673</guid></item><item><title><![CDATA[New comment by easygenes in "GLM-5.2: The Most Powerful Open Model yet and the Brutal Reality of Running It"]]></title><description><![CDATA[
<p>Article reads as though written by someone who doesn't have much experience with deployments like this. Underestimates the memory needed to run with a reasonable amount of context. Misses two other obvious targets:<p><pre><code>  1) 4x DGX Spark (or equivalent other GB10 boxes) with a switch (MikroTik CRS504 or CRS804) and TP=4.
  2) 4x RTX PRO 6000 box. Probably the most practical for cost/perf if you want on-prem as an individual.
</code></pre>
Both would be best to run a 2-bit quant so everything can stay resident (article claims you could run a 4-bit quant with 4x RTX 6000 Ada, and while technically true it would mean a lot of the weights are streaming from DRAM, so it would be slow and impractical. You would need 8x RTX PRO 6000 to run 4 bit at a good speed).<p>This model quantizes unusually well: <a href="https://unsloth.ai/docs/models/glm-5.2#quantization-analysis">https://unsloth.ai/docs/models/glm-5.2#quantization-analysis</a></p>
]]></description><pubDate>Fri, 19 Jun 2026 02:39:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=48594256</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48594256</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48594256</guid></item><item><title><![CDATA[New comment by easygenes in "The Korean telecom giant at the center of Anthropic's Mythos controversy"]]></title><description><![CDATA[
<p>The Wired headline reframes the issue in a way that’s misleading. SK Telecom was a previously resolved issue (as in prior to Fable launch).<p>It may have been a contributing factor, but the crux of the shutdown was the industry reporting of Fable jailbreaks (reportedly spearheaded by Amazon CEO Andy Jassy). The more interesting and honest angle is that the industry which has taken the seriousness of Glasswing at face value felt blindsided by Fable release and totally exposed by the residual risk, when they know they still have a months-long bugfixing backlog exposed by Glasswing and are desperate to buy more time.<p>This misleading looks deliberate on Wired’s part, to appear as though they’re getting a scoop when they’re really just being dishonest. Shameful.</p>
]]></description><pubDate>Thu, 18 Jun 2026 19:59:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=48590700</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48590700</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48590700</guid></item><item><title><![CDATA[New comment by easygenes in "OpenAI Losses Increased Nearly 8X in 2025, with Spending Hitting $34B"]]></title><description><![CDATA[
<p>This headline is not what I would read from this. The numbers are more favorable than the general tone of rumors, and point towards the expected shape of a fast-growing R&D heavy business.</p>
]]></description><pubDate>Thu, 18 Jun 2026 00:21:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48578878</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48578878</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48578878</guid></item><item><title><![CDATA[New comment by easygenes in "GLM 5.2 Is Out"]]></title><description><![CDATA[
<p>Announcement from the founder of Z.ai:<p>“ GLM-5.2 is Fully Open, Frontier Intelligence Belongs to Everyone<p>Today, the sudden restriction of certain frontier models is deeply regrettable. At a time when access to frontier models is abruptly cut off for non-technical reasons, we are even more convinced of one thing: science should be global.<p>The path to AGI (Artificial General Intelligence) must never be enclosed by high walls. We have always believed that AGI should be the cornerstone for all of humanity to collaboratively explore the boundaries of intelligence and solve complex challenges, rather than a privilege monopolized by a few rules and subject to revocation at any moment. In the face of external blockades and restrictions, our attitude is one of radical openness. Frontier intelligence must remain open-source, accessible, and buildable, serving every dedicated developer.<p>GLM-5.2 is Zhipu's most capable open-source model to date. It not only supports a truly usable 1M context window but also maintains a continuous lead in the independent completion of long-horizon tasks, providing solid foundational support for building complex agent applications. It also continues to be our main engine for creating the strongest domestic coding model.<p>Tonight at 5:21—at this special moment—GLM-5.2 will officially be available to all GLM Coding Plan users (including Lite / Pro / Max). The API will also go live next week.<p>A step closer to frontier intelligence for everyone.
The future of AI is open, and it is for the people.
ModelKey: GLM-5.2”<p><a href="https://x.com/jietang/status/2065784751345287314" rel="nofollow">https://x.com/jietang/status/2065784751345287314</a></p>
]]></description><pubDate>Sat, 13 Jun 2026 20:32:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=48521149</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48521149</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48521149</guid></item><item><title><![CDATA[New comment by easygenes in "GLM 5.2 Is Out"]]></title><description><![CDATA[
<p>This release was rushed to hang on the coattails of the Mythos drama (“hey, sorry you can’t use Fable, but try us while you wait this weekend!”) I think they planned to release next week, hence benchmarks not all being ready yet.</p>
]]></description><pubDate>Sat, 13 Jun 2026 19:04:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=48520363</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48520363</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48520363</guid></item><item><title><![CDATA[New comment by easygenes in "Nvidia is proposing a beast of a CPU system for Windows PCs"]]></title><description><![CDATA[
<p>That happened a year ago when these shipped as the DGX Spark with only Linux pre installed.</p>
]]></description><pubDate>Sun, 07 Jun 2026 11:47:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48433934</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48433934</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48433934</guid></item><item><title><![CDATA[New comment by easygenes in "Nvidia is proposing a beast of a CPU system for Windows PCs"]]></title><description><![CDATA[
<p>Mostly a strategy move to protect the CUDA moat… Apple would take over mobile inference in a clean sweep without competition.</p>
]]></description><pubDate>Sun, 07 Jun 2026 11:45:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48433923</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48433923</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48433923</guid></item><item><title><![CDATA[New comment by easygenes in "Nvidia is proposing a beast of a CPU system for Windows PCs"]]></title><description><![CDATA[
<p>This is the same chip and same memory. Only difference is it is going in a laptop, so will be more thermally limited.</p>
]]></description><pubDate>Sun, 07 Jun 2026 11:12:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48433733</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48433733</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48433733</guid></item><item><title><![CDATA[New comment by easygenes in "Uber's $1,500/month AI limit is a useful signal for AI tool pricing"]]></title><description><![CDATA[
<p>If I were paying API rates this year, I would have already burned through $20k in tokens. Looking forward to the costs of this level of capability coming down.</p>
]]></description><pubDate>Wed, 03 Jun 2026 23:58:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48391828</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48391828</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48391828</guid></item><item><title><![CDATA[New comment by easygenes in "Gemma 4 12B: A unified, encoder-free multimodal model"]]></title><description><![CDATA[
<p>I have now also tried it on this scatter plot: <a href="https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FtRVN97QDzO0Pq7SscC7x%2Fgemma%20426b%20bench.png?alt=media&token=80b4da76-efe9-4554-8e31-cca6494d456c" rel="nofollow">https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-p...</a><p>Similarly, the 26B A4B Gemma 4 and the 35B A3B Qwen 3.6 identify it clearly, give me the title and trends analysis fairly accurately. While this 12B spits out gobbledygook about it having something to do with hard-drive capacity. It's like it can barely see, gets the very broad strokes (knows it's looking at some kind of chart), but can't identify any details clearly.</p>
]]></description><pubDate>Wed, 03 Jun 2026 23:45:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=48391687</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48391687</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48391687</guid></item><item><title><![CDATA[New comment by easygenes in "Gemma 4 12B: A unified, encoder-free multimodal model"]]></title><description><![CDATA[
<p>They haven't made one for this new model, but Unsloth has a comprehensive quant KLD map of Gemma 4 26B A4B here: <a href="https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FxhOjnexMCB3dmuQFQ2Zq%2Fuploads%2FtRVN97QDzO0Pq7SscC7x%2Fgemma%20426b%20bench.png?alt=media&token=80b4da76-efe9-4554-8e31-cca6494d456c" rel="nofollow">https://3215535692-files.gitbook.io/~/files/v0/b/gitbook-x-p...</a></p>
]]></description><pubDate>Wed, 03 Jun 2026 23:36:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=48391593</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48391593</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48391593</guid></item><item><title><![CDATA[New comment by easygenes in "Gemma 4 12B: A unified, encoder-free multimodal model"]]></title><description><![CDATA[
<p>I want to like the vision capabilities of the model. However, when I gave it an image which Gemma 26B A4B and Qwen 3.6 35B A3B has no problem correctly describing in detail, including identifying the Taj Mahal in the background it utterly failed. Its sense of the image was that it was a "distorted wide panorama" and even when I asked directly if it was the Taj Mahal it said no. The reference models saw it correctly as a normal square image taken from a fairly rectilinear lens (iPhone main camera).</p>
]]></description><pubDate>Wed, 03 Jun 2026 23:24:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48391488</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48391488</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48391488</guid></item><item><title><![CDATA[New comment by easygenes in "MAI-Code-1-Flash"]]></title><description><![CDATA[
<p>Have you run it through DeepSWE? I understand that's probably a high ask for this class of model, but would be interesting to see regardless.<p>Even if it can't fully pass much, there are so many tests against most of the scenarios that you can get a fairly rich report beyond the pass@1 stat. See e.g. this DeepSWE report against the Minimax M3 model: <a href="https://entrpi.github.io/misc/deep-swe-minimax-m3/" rel="nofollow">https://entrpi.github.io/misc/deep-swe-minimax-m3/</a></p>
]]></description><pubDate>Wed, 03 Jun 2026 07:53:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=48381166</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48381166</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48381166</guid></item><item><title><![CDATA[New comment by easygenes in "MAI-Code-1-Flash"]]></title><description><![CDATA[
<p>While I agree directionally, I'll caveat that "cost per token" != "cost per task". In the case of Qwen3.6 it tends to think 1.6x more than Haiku, so the cost of Haiku on the same tasks tends to only be about double. More detail from comparing their Artificial Analysis metrics:<p><pre><code>  Qwen3.6-35B-A3B   vs   Claude Haiku 4.5
    reasoning mode · AA Intelligence Index v4.0
  
  46.0 ┤   ↖ better — cheaper · smarter · faster
       │
       │
  44.0 ┤     ╭─────╮
       │     │  ●  │ Qwen3.6-35B-A3B
       │     ╰─────╯
  42.0 ┤
       │
       │
  40.0 ┤
       │
       │
  38.0 ┤                                       ╭───╮
       │                      Claude Haiku 4.5 │ ○ │
       │                                       ╰───╯
  36.0 ┤
       └┬─────────┬─────────┬─────────┬─────────┬────────┬
        $200    $300      $400      $500      $600    $700
  
    x → cost to run the index (USD)        lower is better
    y → AA intelligence index              higher is better
  
    bubble area = output speed (tokens / sec)
          ╭─────╮                  ╭───╮
          │  ●  │ Qwen ~196 t/s    │ ○ │ Haiku ~93 t/s
          ╰─────╯                  ╰───╯
  
    ┌─────────────────────┬──────────┬──────────┬───────────┐
    │ model               │ AA index │ run cost │ out speed │
    ├─────────────────────┼──────────┼──────────┼───────────┤
    │ Qwen3.6-35B-A3B    ●│   43.5   │   $280   │  196 t/s  │
    │ Claude Haiku 4.5   ○│   37.1   │   $620   │   93 t/s  │
    └─────────────────────┴──────────┴──────────┴───────────┘


    COST PER TOKEN   ≠   COST PER TASK  
    output tokens per index run:
       Haiku 4.5    87.3M   (79.3M reasoning + 8.0M answer)
       Qwen3.6     143.2M   (131.7M reasoning + 11.5M answer)
       → Qwen emits 1.64× more output
  
    ── output speed (tokens / sec) ──────────  raw rate · higher = faster
       Qwen3.6     100%   ~196 t/s
       Haiku 4.5   ~47%   ~93 t/s
                                                  → Qwen ~2.1× faster per token
  
          ╎   1.64× more tokens  <  2.1× faster rate
          ▼
  
    ── solution speed (per finished answer) ──  higher = faster
       Qwen3.6     100%
       Haiku 4.5   ~78%
                                                  → Qwen ~1.3× FASTER to a solution
  
    SCORECARD
                            intelligence    cost / task     speed to solution
     Qwen3.6-35B-A3B        43.5            $280            ~1.3× faster 
     Claude Haiku 4.5       37.1            $620            (slower)
  
     → Qwen wins all three. The reasoning blow-up (1.64×) is smaller than
       the raw-speed edge (2.1×), so Qwen stays ahead per task.</code></pre></p>
]]></description><pubDate>Wed, 03 Jun 2026 03:57:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48379719</link><dc:creator>easygenes</dc:creator><comments>https://news.ycombinator.com/item?id=48379719</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48379719</guid></item></channel></rss>