<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: neosat</title><link>https://news.ycombinator.com/user?id=neosat</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 29 Sep 2026 04:17:54 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=neosat" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by neosat in "Ember-1"]]></title><description><![CDATA[
<p>Lots of use cases!
I've personally used it for the following:<p>1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.<p>2. e-commerce catalog classification
3. quick search using anything as context and query mapping to a pre-defined set.</p>
]]></description><pubDate>Sun, 27 Sep 2026 18:59:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49869728</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49869728</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49869728</guid></item><item><title><![CDATA[New comment by neosat in "A misalignment of AI in mathematics"]]></title><description><![CDATA[
<p>>> "are these companies interested in developing research"
judging from the money, resources spent and the value they derive from this the answer is very definitively yes.<p>What makes you think these companies (and I'm not a fan of all their motives) are not interested in developing research? The motives may be self-serving, but it is undoubtedly and objectively accelerating research.</p>
]]></description><pubDate>Sat, 12 Sep 2026 04:41:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49668888</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49668888</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49668888</guid></item><item><title><![CDATA[New comment by neosat in "A misalignment of AI in mathematics"]]></title><description><![CDATA[
<p>You can couch it different words but the basic shape of all these is the same, whether it was voice artists earlier, IT outsourcing, or now disciplines like mathematics. A small set of people (relatively) who were the primary source of getting something done, suddenly find that technology has made it accessible for others to do what they specialized in.  It's a tough pill to swallow and it is natural to not be comfortable with this for most people.<p>However, as it has happened in the past, once there is a technological wedge, technological advances will move forward, whether some community likes it or not, and human ingenuity will find ways such that the benefit is greater than the risk.</p>
]]></description><pubDate>Sat, 12 Sep 2026 03:56:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49668615</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49668615</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49668615</guid></item><item><title><![CDATA[New comment by neosat in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>Can you or someone else from A\ comment on whether the conversation style is coming to Opus 5 or a future 5.1 asap as well? Currently it seems the model has been made unusable by the way it 'speaks' and there is a clear solution where it can speak better but nothing has been done about the flagship model on Pro plans. I've literally had to work on Opus 4.8 which does not have this problem and speaks fine.</p>
]]></description><pubDate>Tue, 01 Sep 2026 20:25:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49527649</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49527649</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49527649</guid></item><item><title><![CDATA[New comment by neosat in "You Probably Don't Get Why Stripe Bought OpenRouter"]]></title><description><![CDATA[
<p>Sorry but the authors of this article probably don't fully 'get it' (or at least cannot explain it well) either if this is their conclusion.<p>"In other words, Stripe did not buy a router. It bought a strategic frontier AI systems security and alignment asset. We believe combining this asset with Stripe’s existing infrastructure strengthens the ecosystem’s independent alignment capabilities in ways that would be difficult for any single lab to accomplish by itself. As such, this acquisition is net positive for the health of the independent frontier ecosystem."<p>Contrast with the clarity reflected in the letter Stripe reportedly sent to its investors.<p>"We see capital and intelligence are becoming the two digital flows undergirding every business ..... going forward every developer will also need a straightforward and reliable way to manage their intelligence pipeline"  
So basically Stripe already enables companies to manage capital. It now also hopes, with the acquisition, to jump to the forefront of managing intelligence / tokens for them.  By bundling two core capabilities for AI forward companies they increase their offering attractiveness to businesses.</p>
]]></description><pubDate>Wed, 19 Aug 2026 20:22:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49366713</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49366713</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49366713</guid></item><item><title><![CDATA[New comment by neosat in "DeepSeek V4 Pro 0813"]]></title><description><![CDATA[
<p>As the above two comments mentioned this is not true in practice due to batch effects (you can read about some interesting work published by Thinking Machines on this), as well as calculation drift that happens across computations esp. now with inference optimization becoming common.</p>
]]></description><pubDate>Thu, 13 Aug 2026 00:30:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49280381</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49280381</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49280381</guid></item><item><title><![CDATA[New comment by neosat in "Qwen3.8-2.4T"]]></title><description><![CDATA[
<p>This may not 'quite a bit behind' those at all. If you look at the benchmark numbers they are very comparable to Fable, but beyond a certain point the benchmark numbers don't tell you much. Opus #5 beats Fable on some benchmarks but given similar cost almost everyone who has used those two models will prefer to use Fable.<p>At this price range $0.87per 1M they will get a lot of usage of people trying it out. Given the benchmark numbers, for many people and many use cases this will become their primary driver. There are people and use cases where Fable, Sol will work better but those are likely not the target of DeepSeek anyway.<p>In terms of performance and price pareto curve I don't think any model can beat this today (though openAI is doing some exciting recent work in efficiency) - which is a remarkable feat for the DeepSeek team.<p>Either way, what a time for consumers of these models :)</p>
]]></description><pubDate>Wed, 12 Aug 2026 16:43:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49275206</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49275206</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49275206</guid></item><item><title><![CDATA[New comment by neosat in "Elevated errors on Claude Opus 5"]]></title><description><![CDATA[
<p>Yes I had a similar experience with Opus 5. It is very token efficient, fast, and gets reasonable part of the work right but makes a LOT of mistakes. In a month+ use of Fable completed each task without ANY errors. Opus could not complete a single of ~5 tasks without some issue or the other - either not getting it fully right or actually introducing regressions. To their credit it was able to catch regressions and fix competently.  It seems like a pre Opus 4.6 model in terms of reliability with a lot more power and spiky intelligence. When it gets things right it's powerful and efficient but without reliability I had to 'downgrade' to Opus 4.8 forcibly (since it was not a default option on claude code). I really miss Fable on the pro plan and will likely churn to K3.</p>
]]></description><pubDate>Mon, 27 Jul 2026 18:14:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49073488</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=49073488</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49073488</guid></item><item><title><![CDATA[New comment by neosat in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>Definitely not my experience. Fable is better but I'd prefer K3 to Opus based my experience with both.</p>
]]></description><pubDate>Sat, 18 Jul 2026 23:05:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=48963357</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48963357</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48963357</guid></item><item><title><![CDATA[New comment by neosat in "Previewing GPT‑5.6 Sol: a next-generation model"]]></title><description><![CDATA[
<p>Good observations. There's definitely a trend in pricing increasing but also balanced by innovations and availability of other models (both open and closed) emerging as alternatives. It's natural for the labs to explore how much they can push pricing, and for competitors to explore how they can treat that margin as their opportunity to grow their business.<p>Eventually the pricing should be more stable.</p>
]]></description><pubDate>Fri, 26 Jun 2026 17:25:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48689291</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48689291</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48689291</guid></item><item><title><![CDATA[New comment by neosat in "GLM-5.2 is a step change for open agents"]]></title><description><![CDATA[
<p>I've been using GLM 5.2 recently (company hosted, for non-coding tasks) and it's been strong and reliable. There are areas where GPT 5.5 and Opus 4.x still feel marginally better but only marginally. For most tasks if GLM 5.2 is the only model I have to use I'm productive and happy. This was not true before GLM 5.2.  No doubt in my mind that the gap is closing quickly and for most tasks that are not very specialized open models will be usably on par on flagship closed models and have an edge factoring in cost.<p>For coding I still use 5.5 w/ Codex and prefer that to other models + harness combinations.</p>
]]></description><pubDate>Wed, 24 Jun 2026 23:51:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=48666962</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48666962</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48666962</guid></item><item><title><![CDATA[New comment by neosat in "Identity verification on Claude"]]></title><description><![CDATA[
<p>Anthropic is really on a tricky path here. When you have had runaway success due to a hit it is easy to believe that it is the natural way of things. However, that happened due to unique convergence of tech paradigm shift, the competitive landscape,  and how they were positioned to capture that value through claude code.<p>They somehow conflate their value with 'safety'. While it's an admirable internal quality for the company to have, their treatment of their user base (developers, users) has been bordering on indifference and their stance bordering on arrogance.<p>As competition heats up, there is a very real chance of them shooting themselves in the foot with friction such as this (to be fair not completely in their control but also they had their share of responsibility that led to this)</p>
]]></description><pubDate>Sun, 21 Jun 2026 20:57:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=48622532</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48622532</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48622532</guid></item><item><title><![CDATA[New comment by neosat in "Gemma 4 12B: A unified, encoder-free multimodal model"]]></title><description><![CDATA[
<p>Agree. Audio has strongly temporal so there is almost certainly some positional encoding one way or another.</p>
]]></description><pubDate>Wed, 03 Jun 2026 17:39:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48387096</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48387096</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48387096</guid></item><item><title><![CDATA[New comment by neosat in "Anthropic surpasses OpenAI to become most valuable AI startup"]]></title><description><![CDATA[
<p>You need to see the response in light of the original discussion. Referencing here for clarity since I should have included it in the first place: "We used the claude code and codex harness and I implemented some prs they needed with gpt5.5 and opus4.7 and asked them to identify which came from which only from the code."<p>So the same person, was using similarly competitive tools, and showing that the output was hard to discern (indirectly the implication was also that implementation was fairly trivial in both of those). A better analogy would not be different process and widely different tools but for example two power drills. Sure, folks could still prefer one over the other, but that's a different claim that saying X is objectively better than Y when both are directly competing on very similar dimensions.<p>Assuming you meant Claude code: I'd love to learn more about "Codex and Claude are very different" because maybe I'm assuming just based on my use case where I use both of them interchangeably for the same thing (coding web and mobile apps)</p>
]]></description><pubDate>Sat, 30 May 2026 16:06:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48337709</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48337709</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48337709</guid></item><item><title><![CDATA[New comment by neosat in "Anthropic surpasses OpenAI to become most valuable AI startup"]]></title><description><![CDATA[
<p>That's a fair callout and I agree my statement was too general in just mentioning 'output', as you correctly pointed out. To define 'better' you would indeed need to agree on the dimensions you would evaluate candidates against.<p>I think a more appropriate rephrasing would be 'You cannot simply make a claim that (model + harness) X is better than Y, but then have no discernible difference on dimensions you care about'.    In the case of latest of claude code vs codex with gpt 5.5) both are similar enough in the dimensions people will care about in evaluating (vs. differing wildly in cost or time taken).</p>
]]></description><pubDate>Sat, 30 May 2026 14:58:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48336966</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48336966</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48336966</guid></item><item><title><![CDATA[New comment by neosat in "Anthropic surpasses OpenAI to become most valuable AI startup"]]></title><description><![CDATA[
<p>Your argument is fine but different from the claim the OP is making. You cannot simply make a claim that (model + harness) X is better than Y, but then have no discernible difference in the output.  Subjectively, people might still prefer one over due to anything from design to marketing, but that's very different from the claim that X is better than Y for coding (see: "A colleague was convinced Claude is better").   Basically, I prefer Claude is a different claim than Claude is better and the latter has a higher bar of proof.</p>
]]></description><pubDate>Sat, 30 May 2026 14:39:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=48336734</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48336734</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48336734</guid></item><item><title><![CDATA[New comment by neosat in "Why Japanese companies do so many different things"]]></title><description><![CDATA[
<p>Exactly, I was confused too. The authors clearly mention what the parent comment talks about, albeit towards the end of the article, that the 'J' bundle meant that these firms were not set up for success once they 'caught up' and were required to innovate not just process but from the ground up to envision new categories (e.g. iPhone).</p>
]]></description><pubDate>Fri, 22 May 2026 20:07:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48240909</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48240909</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48240909</guid></item><item><title><![CDATA[New comment by neosat in "SpaceX S-1"]]></title><description><![CDATA[
<p>Revenue is not the right metric when you compare space trips to trips inside a city. The more relevant numbers are EBITDA, Operating cash flow, Profits.</p>
]]></description><pubDate>Thu, 21 May 2026 00:21:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48216186</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48216186</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48216186</guid></item><item><title><![CDATA[New comment by neosat in "SpaceX S-1"]]></title><description><![CDATA[
<p>has anyone done the math on:
1. cost to build out and run the data centers
2. cost of compute (hardware and energy)
3. depreciation of legacy GPU and thus value at the end of 3 years.<p>And then compare the $45B revenue from Anthropic to see if it's mostly break even or if one of Anthropic/SpaceX came out ahead on the contract.</p>
]]></description><pubDate>Wed, 20 May 2026 21:56:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=48214760</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48214760</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48214760</guid></item><item><title><![CDATA[New comment by neosat in "Show HN: Lance – image/video generation and understanding in one model"]]></title><description><![CDATA[
<p>That's true, I should have mentioned active. Actual params are closer to 12B-14B likely, given the 40GB VRAM usage.</p>
]]></description><pubDate>Wed, 20 May 2026 20:38:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48213773</link><dc:creator>neosat</dc:creator><comments>https://news.ycombinator.com/item?id=48213773</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48213773</guid></item></channel></rss>