<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: x313</title><link>https://news.ycombinator.com/user?id=x313</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 21 Jul 2026 19:55:52 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=x313" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by x313 in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a guaranteed-correct implementation. Similar business model as open-source SaaS companies.<p>Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.</p>
]]></description><pubDate>Mon, 20 Jul 2026 17:41:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48982131</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48982131</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48982131</guid></item><item><title><![CDATA[New comment by x313 in "China’s open-weights AI strategy is winning"]]></title><description><![CDATA[
<p>Similar story here. DS models are absurdly good value for mid-end tasks. I've found DSv4 Flash to be ~10% the cost of GPT-5.4-mini/Claude Haiku at similar performance.<p>We used to pay OpenAI >1m$/month for fraud classification, NER, etc. Sadly the US companies no longer care about non-coding-agent uses.<p>I imagine uptake will continue to increase as the corporate infra improves. Right now it's still bad - for example, AWS Bedrock is awful, models are months late and implemented with basic errors. Google Vertex is even worse. Finding a decent provider is the hardest part.</p>
]]></description><pubDate>Mon, 20 Jul 2026 17:22:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48981856</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48981856</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48981856</guid></item><item><title><![CDATA[New comment by x313 in "Evidence of inconsistencies in evaluation process and selection of winners"]]></title><description><![CDATA[
<p>With all due respect, there's zero chance that humans with relevant knowledge scored these themselves. Reading through the winners, every single one is classic vibe-research, with the usual pure-LLM-research patterns:<p>- Grand claims backed by no evidence<p>- Core designs that make zero sense<p>- Pointless graphs that show nothing of interest<p>- Yet endless robustness checks on minor methodological assumptions (especially confidence intervals and t-tests)<p>For example, on the winning entry, not only is the graph completely wrong (as mentioned by the OP), but the interpretation would be nonsensical even if it was (implying bigger models "get more RL"?). And their own results even show the core dataset is worthless, because all their metrics are near-perfectly correlated. There's no way a serious human reader trying to evaluate "is this benchmark useful" would ever miss this.<p>I don't mean to pick on them - all the winning entries seem like there was no human effort put into them. And again, there's no way a human who actually attempted to read and understand these would ever think these are good by even the most minimal of standards.<p>Kaggle is legitimately a really awesome website, as someone who's competed before and won a few contests pre-LLMs. Stuff like this winning devalues the entire product and makes it look like a joke. If almost all entries look like this now, it'd be better to allow for the possibility of no winner.</p>
]]></description><pubDate>Sat, 18 Jul 2026 02:11:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48954484</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48954484</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48954484</guid></item><item><title><![CDATA[New comment by x313 in "Evidence of inconsistencies in evaluation process and selection of winners"]]></title><description><![CDATA[
<p>This is jaw-droppingly lazy slop. The authors really didn't put in even an ounce of thought or effort.</p>
]]></description><pubDate>Fri, 17 Jul 2026 16:45:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48949410</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48949410</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48949410</guid></item><item><title><![CDATA[New comment by x313 in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>If there was some grand strategy for all Chinese labs, surely it'd have leaked by now. I think its more likely that:<p>- Companies can still make money from commodities<p>- Chinese labs only have 5-10% the valuation of OpenAI/Anthropic, so massive monopoly profits aren't necessary. Profit expectations for tech companies in China are really low in general, complete opposite of the US.<p>- Open weighting is a great way to get talent/attention/reputation</p>
]]></description><pubDate>Thu, 16 Jul 2026 22:24:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48941078</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48941078</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48941078</guid></item><item><title><![CDATA[New comment by x313 in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>Strictly dominates both Sonnet 5 and Opus 4.8 in both cost and performance:<p><a href="https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-claude-opus-4-8" rel="nofollow">https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...</a><p><a href="https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-claude-sonnet-5" rel="nofollow">https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-...</a></p>
]]></description><pubDate>Thu, 16 Jul 2026 21:03:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=48940216</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48940216</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48940216</guid></item><item><title><![CDATA[New comment by x313 in "AI Hiring Tools Yield Racial Bias and Systemic Rejection; 26% Black & 15% Asian"]]></title><description><![CDATA[
<p>This study only looks at one specific vendor algorithmn (a job assesment given by a company called pymetrics)</p>
]]></description><pubDate>Tue, 23 Jun 2026 20:23:45 +0000</pubDate><link>https://news.ycombinator.com/item?id=48650836</link><dc:creator>x313</dc:creator><comments>https://news.ycombinator.com/item?id=48650836</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48650836</guid></item></channel></rss>