<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: ponyous</title><link>https://news.ycombinator.com/user?id=ponyous</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 15 Sep 2026 09:07:06 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=ponyous" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by ponyous in "Steam Frame starts at $1059"]]></title><description><![CDATA[
<p>You mean people blind in one eye? I think it still works, but obviously depth perception will not be there</p>
]]></description><pubDate>Tue, 15 Sep 2026 07:51:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49709171</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49709171</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49709171</guid></item><item><title><![CDATA[New comment by ponyous in "Ask HN: What are you working on? (September 2026)"]]></title><description><![CDATA[
<p>Surprisingly Gemini models. Spatial understanding seems to be the best for the price and speed.</p>
]]></description><pubDate>Mon, 14 Sep 2026 06:54:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=49692903</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49692903</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49692903</guid></item><item><title><![CDATA[New comment by ponyous in "Ask HN: What are you working on? (September 2026)"]]></title><description><![CDATA[
<p><a href="https://grandpacad.com" rel="nofollow">https://grandpacad.com</a> - AI modeling for 3D printing
Been at it for about year and a half.<p>Really exciting stuff is happening literally every month, because underlying models are getting better and better. When I started it was pretty basic: “make a cube with a hole through it”. Now it’s at the point “make a raspberry pi 4 case” and the agent searches, builds, verifies…<p>What surprised me in this process is how little meaning AI benchmarks have. Pareto frontier for my use case looks completely different than any other benchmark portrays.</p>
]]></description><pubDate>Sun, 13 Sep 2026 18:00:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49686727</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49686727</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49686727</guid></item><item><title><![CDATA[New comment by ponyous in "DeepSeek v4.1 Flash"]]></title><description><![CDATA[
<p>> I have a feeling that the price of 1 million tokens transmitted over the internet is more expensive than cache hit.<p>And this kinda makes sense. What is cheaper few KB of disk space or internet bandwidth?</p>
]]></description><pubDate>Thu, 10 Sep 2026 12:23:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49642585</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49642585</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49642585</guid></item><item><title><![CDATA[New comment by ponyous in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>We went for 16% intelligence bump according to artificial analysis for +82% of the cost. Interesting.<p>Comparing 4.8 Opus with Fable 5.1</p>
]]></description><pubDate>Wed, 02 Sep 2026 09:32:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49533889</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49533889</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49533889</guid></item><item><title><![CDATA[New comment by ponyous in "Gemini 3.7 Flash"]]></title><description><![CDATA[
<p>On my benchmark where AIs generate ~20 different 3D models about 1/2 the time of Opus and 1/3 of the time of Kimi K3 and 2/3 of time of sonnet.</p>
]]></description><pubDate>Thu, 13 Aug 2026 18:16:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49289922</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49289922</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49289922</guid></item><item><title><![CDATA[New comment by ponyous in "Ask HN: What are you working on? (August 2026)"]]></title><description><![CDATA[
<p>Yeah you can absolutely do multiple parts, although the mating features are still a bit rough you can do a lot already.</p>
]]></description><pubDate>Mon, 10 Aug 2026 09:26:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49241338</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49241338</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49241338</guid></item><item><title><![CDATA[New comment by ponyous in "Ask HN: What are you working on? (August 2026)"]]></title><description><![CDATA[
<p>GrandpaCAD - AI 3D modeling software focused on simplicity. Made it so even my grandpa could model.
He’s been asking me for years when will I teach him how to 3D model. I tried, we failed and then I seen him use ChatGPT so I knew there was a better way than traditional CAD tools.<p>Recently we also got European funding and the project got some traction. Very exciting times ahead.<p><a href="https://grandpacad.com" rel="nofollow">https://grandpacad.com</a></p>
]]></description><pubDate>Sun, 09 Aug 2026 19:19:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49234709</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49234709</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49234709</guid></item><item><title><![CDATA[New comment by ponyous in "DeepSeek V4 Flash 0731"]]></title><description><![CDATA[
<p>You are right, relatively to other llm providers this is not slow. But if you think what is possible when you have 1000t/s a sec you might find it slow.</p>
]]></description><pubDate>Fri, 07 Aug 2026 21:50:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49216642</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49216642</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49216642</guid></item><item><title><![CDATA[New comment by ponyous in "Modern email can be built from borrowed parts"]]></title><description><![CDATA[
<p>Spark email client supports this. It’s great</p>
]]></description><pubDate>Mon, 27 Jul 2026 17:06:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49072566</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49072566</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49072566</guid></item><item><title><![CDATA[New comment by ponyous in "Are AI labs pelicanmaxxing?"]]></title><description><![CDATA[
<p>I agree with the conclusion and am happy to see this blog post, but this killed a bit of credibility for me:<p>> Using a single LLM judge for scoring. Every score here comes from one model, GPT-5.6 Luna, looking at one image at a time. I didn’t do much alignment and didn’t check how often it agrees with itself on a re-run.<p>Having used a similar setup (with previous gen LLMs) to evaluate the 3D models that my product[0] generates, it turned out there was no correlation at all. LLM judgments were very much random and I assume judging SVGs is not that far from judging 3D models. I guess I have to re-test this with current gen.<p>[0]: <a href="https://grandpacad.com" rel="nofollow">https://grandpacad.com</a></p>
]]></description><pubDate>Thu, 23 Jul 2026 12:54:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49020810</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=49020810</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49020810</guid></item><item><title><![CDATA[New comment by ponyous in "Hy3"]]></title><description><![CDATA[
<p>Interesting way to show off a model last on every benchmark. Not sure any other lab is doing this</p>
]]></description><pubDate>Fri, 10 Jul 2026 10:36:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48858146</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48858146</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48858146</guid></item><item><title><![CDATA[New comment by ponyous in "Ask HN: Has anyone had success with SBIR grants and what is the process like?"]]></title><description><![CDATA[
<p>As someone who just had success with equivalent system in EU my recommendation is to get someone who's done it before to do it for you. I hired an agency. They took 8% fee, which is pretty low, usually it's between 10% and 15%.</p>
]]></description><pubDate>Thu, 18 Jun 2026 10:08:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=48583207</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48583207</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48583207</guid></item><item><title><![CDATA[New comment by ponyous in "GLM-5.2 is the new leading open weights model on Artificial Analysis"]]></title><description><![CDATA[
<p>Absolutely. Running it now, will update this comment in about 30 mins.<p>Edit: Surprisingly very good results with 3.0 flash with high thinking.<p>Cost: $0.06<p>Duration: 3.22 min<p>Code Errors: 1.3 per attempts (meaning on average it had to retry 1.3 times)<p>Adherence was on par with 3.5 flash Low thinking</p>
]]></description><pubDate>Wed, 17 Jun 2026 14:28:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48571043</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48571043</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48571043</guid></item><item><title><![CDATA[New comment by ponyous in "GLM-5.2 is the new leading open weights model on Artificial Analysis"]]></title><description><![CDATA[
<p>I don't have the eval results live yet, so I cannot share them yet.<p>I was benchmarking using a soon to be released new version of my AI CAD modeling software[0].
It's basically an agent that has access to tools that can execute build123d scripts, get sculpted models, blender to combine sculpts + parametric models, tools to inspect the model (visually and with code), search datasheets, ...<p>I tried what you recommend a while ago (asking an AI to evaluate using different angles) and the AI evaluations were extremely bad - barely any correlation to what I scored. Things have gotten better, but I don't trust it enough yet.<p>Here is how I score adherence (and how AI did as well, but I tried methods where it would just give back a boolean "pass" or not):<p><pre><code>    <0.2 → Poor – Misses core intent; largely irrelevant or incorrect.
    <0.4 → Weak – Partially relevant; significant omissions or errors.
    <0.6 → Fair – Covers main points but lacks completeness or precision.
    <0.8 → Good – Mostly accurate; minor gaps or deviations.
    <=1.0 → Excellent – Fully aligned; precise, comprehensive, and faithful to intent.
</code></pre>
Here is the scenario list (prompts are much more detailed):<p><pre><code>    dragon-bottle-stopper
    editing-param-mid-conv
    editing-parametric-enclosure
    editing-swap-material-param
    editing-text-edit-cube
    multi-turn-bird-house
    multi-turn-dice-tower
    multi-turn-modular-planter
    multi-turn-phone-stand
    multi-turn-shelf
    one-shot-bookend
    one-shot-cable-clip
    one-shot-chess-queen
    one-shot-coaster
    one-shot-coffee-cup
    one-shot-dog-tag
    one-shot-dragon-figurine
    one-shot-hex-bracket
    one-shot-keychain-fob
    one-shot-low-poly-tree
    one-shot-pegboard-hook
    one-shot-pi4-case
    one-shot-threaded-jar


</code></pre>
[0]: <a href="https://grandpacad.com" rel="nofollow">https://grandpacad.com</a></p>
]]></description><pubDate>Wed, 17 Jun 2026 14:26:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=48571025</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48571025</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48571025</guid></item><item><title><![CDATA[New comment by ponyous in "GLM-5.2 is the new leading open weights model on Artificial Analysis"]]></title><description><![CDATA[
<p>Just ran and scored 63 3d model generations (via code) across high and no reasoning. 3D Modeling benchmark quickly shows spatial, logic and code performance of the model so I think it's a very good indicator of the quality.<p>Here are the results compared to Gemini 3.5 Flash:<p><pre><code>    Model + config          CodeErr/gen   Cost/gen   Median time   Quality
    gemini-3.5-flash, low      0.71        $0.18        68s       baseline
    GLM 5.2, reasoning high    0.61        $0.18       289s         -6.0%
    GLM 5.2, reasoning off     1.52        $0.10       126s        -13.6%

</code></pre>
Although it is cheaper, it is significantly slower, and results are worse overall. Surprisingly - high reasoning produces less code errors than gemini 3.5 flash, but when I actually look at the models they are worse.<p>Edit: I recently ran evals with Kimi 2.7 and MiniMax-M3 and this is clearly open source SOTA model, by far.</p>
]]></description><pubDate>Wed, 17 Jun 2026 13:29:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48570287</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48570287</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48570287</guid></item><item><title><![CDATA[New comment by ponyous in "Claude Fable 5"]]></title><description><![CDATA[
<p>Basically double from Opus 4.8 IIRC</p>
]]></description><pubDate>Tue, 09 Jun 2026 17:38:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=48464504</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48464504</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48464504</guid></item><item><title><![CDATA[New comment by ponyous in "Spanish traders set the standard for GnuCash database design"]]></title><description><![CDATA[
<p>> Surprisingly written by a human :)<p>Article ends with this</p>
]]></description><pubDate>Mon, 08 Jun 2026 13:43:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48445252</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48445252</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48445252</guid></item><item><title><![CDATA[New comment by ponyous in "Bun's unreleased Rust port has 13,365 unsafe blocks"]]></title><description><![CDATA[
<p>Bun is(was?) a lot about performance. How does it compare to zig?</p>
]]></description><pubDate>Fri, 22 May 2026 19:57:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48240776</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48240776</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48240776</guid></item><item><title><![CDATA[New comment by ponyous in "Antigravity 2.0 Tops the OpenSCAD Architectural 3D LLM Benchmark"]]></title><description><![CDATA[
<p>I've run a tons of benchmarks for OpenSCAD for all kinds of models and setups, and what I realised is:<p>- Models are very jagged (might excel in one type of 3d model, but not another)<p>- Gemini models are the least jagged in my experience and have the best image understanding<p>- Gemini models are also the most creative (which may be undesirable if you want precise CAD part)<p>- Overall this benchmark doesn't prove much because one 3d model (and one attempt) is just not enough. I am usually testing on at least a dozen models each generated 3 times, but should really do much more, but it's too pricey for a solo dev.<p>Still, thanks for publishing this. Will be definitely run flash 3.5 soon to see how it performs.</p>
]]></description><pubDate>Fri, 22 May 2026 14:05:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=48236044</link><dc:creator>ponyous</dc:creator><comments>https://news.ycombinator.com/item?id=48236044</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48236044</guid></item></channel></rss>