<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: ricefan1070</title><link>https://news.ycombinator.com/user?id=ricefan1070</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 30 Jul 2026 00:27:31 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=ricefan1070" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by ricefan1070 in "Launch HN: Tokenless (YC S26) – Automatic model switching to save money"]]></title><description><![CDATA[
<p>Does this mean you have to retrain routing rules every time a new model gets released? I imagine since the price/token (or rather the amount of work that can be done per token) does not monotonically increase with new models, that the routing logic has to change all the time</p>
]]></description><pubDate>Wed, 29 Jul 2026 17:21:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49100367</link><dc:creator>ricefan1070</dc:creator><comments>https://news.ycombinator.com/item?id=49100367</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49100367</guid></item><item><title><![CDATA[New comment by ricefan1070 in "[dead]"]]></title><description><![CDATA[
<p>With all the benchmaxxing happening, current evals are oversaturating and become meaningless for model comparison. I think there is only one test that can't be gamed: let them trade in real markets.<p>Markets are self-improving, as models trade better, the market inefficiencies disappear, so trading well gets harder the better you get. It's a never-ending hill to climb.<p>I ran SOTA reasoning LLMs on it. TL;DR: they suck, and reasoning doesn't help. No model came close to a simple static benchmark (experienced human Quant chose parameters);More reasoning ≠ better trading; Fun fact: when losing, LLMs trade less rather than smarter.<p>Why quant trading is a great intelligence test (minimal jargon): a) requires ML research that generalizes OOS (alpha); b) needs robustness to regime shifts and new rules; c) trains long-horizon planning under tradeoffs (how to trade now if AAPL is +2% tomorrow but −5% the day after). And d) it's self-correcting: alphas decay, trading too much gets you adversely selected, and trading well makes the market more efficient — so it never stops being hard.<p>Background: Ex-quants from G-Research (ML algo trading) & TransMarketGroup (Crypto options); + LLM inference kernels at Etched.<p>Please poke holes in the setup - where does "markets as an eval" break down? Especially keen on views from outside quant.</p>
]]></description><pubDate>Wed, 29 Jul 2026 15:23:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49098732</link><dc:creator>ricefan1070</dc:creator><comments>https://news.ycombinator.com/item?id=49098732</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49098732</guid></item></channel></rss>