<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: SilenN</title><link>https://news.ycombinator.com/user?id=SilenN</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 30 Jul 2026 22:51:26 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=SilenN" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Exactly</p>
]]></description><pubDate>Thu, 30 Jul 2026 22:32:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49116710</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49116710</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49116710</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Expensive, in the thousands. We have our own infra in house and are working on bringing these costs down</p>
]]></description><pubDate>Thu, 30 Jul 2026 21:01:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115727</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49115727</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115727</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Technically 0 because 
a) it ingests your already existing traces and does an initial training run
b) in the app we'll have pre-trained routers you can start with that will then learn over time</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:33:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115387</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49115387</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115387</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Fixed formatting which will help with readability. We do routing, distillation, and token compaction.</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:31:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063834</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063834</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063834</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Distill and serve models with frontier quality for half the cost"]]></title><description><![CDATA[
<p>Let me know if you have any questions!</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:31:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063833</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063833</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063833</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Thanks :)</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:31:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063830</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063830</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063830</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Distill and serve models with frontier quality for half the cost"]]></title><description><![CDATA[
<p>Thanks for the heads up, removed mention!</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:30:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063822</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063822</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063822</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Happy to answer any qs.</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:25:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063790</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063790</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063790</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Valid criticism. Happy to answer any qs. We're still working on solidfying results.</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:22:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063778</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063778</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063778</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>It's open source!<p>We do have a platform we'll be launching as well to manage training + serving for you which will require more diligent privacy guarantees.</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:22:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063775</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063775</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063775</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Optimize and serve models with Fable quality at half the cost"]]></title><description><![CDATA[
<p>Open source models.<p>wmo routes requests between frontier models and open source models that continuously train using Tinker. As the smaller models improve, more traffic gets routed to them.<p>Calculating cost is just tokens in/out.</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:21:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063769</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063769</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063769</guid></item><item><title><![CDATA[New comment by SilenN in "Show HN: Distill and serve small models with frontier quality for half the cost"]]></title><description><![CDATA[
<p>That's cool, thanks for sharing!</p>
]]></description><pubDate>Mon, 27 Jul 2026 00:14:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49063728</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063728</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063728</guid></item><item><title><![CDATA[Show HN: Optimize and serve models with Fable quality at half the cost]]></title><description><![CDATA[
<p>Hi HN, we built world-model-optimizer, an open source tool to continually improve a specialized model for an agent.<p>It does this by simulating production tool responses through text world modeling (similar to QwenAgentWorld, summary here <a href="https://x.com/silennai/status/2073887455884058814" rel="nofollow">https://x.com/silennai/status/2073887455884058814</a>).<p>We can then use this to train a router for frontier, OS, and local models (use defaults or pick which ones to optimize against).<p>wmo ingests agent traces, builds the simulation, embeds the traces, runs different models you choose against the simulation scenarios, and then uses a KNN for model selection (similar to <a href="https://arxiv.org/abs/2505.19797" rel="nofollow">https://arxiv.org/abs/2505.19797</a>).<p>- Cache aware: cache is taken into account for the effective price in routing.<p>- Confidence gated: we don't deviate from the best fit model when paired evidence over retrieved neighbors is below 0.5 standard errors or on queries unlike anything in the fit set.<p>- Optimize for cost or quality: train a balanced, cost max, or quality max router.<p>Usage<p>`wmo build` creates the simulation (or add your own benchmark)<p>`wmo optimize` tunes the router<p>`wmo serve` starts the server and can run everything fully locally. The simulation and router can update over time as more agent traces are gathered and new models are added.<p>Router results vs Fable<p>- RouterBench: -66.5% cost, -1.7% performance, -24.7% latency p50. 77.5% of traffic to Sonnet 5, 16.1% Fable 5.<p>- TauBench: -44.5% cost, +6.3% performance, -20% latency. 83% to Opus 5, 17% to Kimi-K2.6 (over K3).<p>- Terminal Bench 2: -64% cost, +8% performance, -50.6% latency. Sonnet 5 is fully along the pareto front. Training a specialized router per task isn't cheap. In sparse data regimes the value can be "here's the best model".<p>We're working on sample effiient continual learning for agent specific models at experientiallabs.ai"</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49063454">https://news.ycombinator.com/item?id=49063454</a></p>
<p>Points: 59</p>
<p># Comments: 27</p>
]]></description><pubDate>Sun, 26 Jul 2026 23:35:15 +0000</pubDate><link>https://github.com/experientiallabs/world-model-optimizer</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=49063454</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49063454</guid></item><item><title><![CDATA[New comment by SilenN in "LLMs as 5x Faster Sandboxes"]]></title><description><![CDATA[
<p>world-model-harness makes it easy to go from agent traces to faithful replication of your production environment where your agents run.<p>Basically, an LLM pretends to be a virtual machine executing instructions. Based on GEPA and Qwen AgentWorld.<p>Just 
- clone,
- `uv sync`
- `uv run wmh build`
and you'll get a wizard that will help you create your own world model from your traces.<p><a href="http://github.com/experientiallabs/world-model-harness" rel="nofollow">http://github.com/experientiallabs/world-model-harness</a></p>
]]></description><pubDate>Tue, 30 Jun 2026 16:30:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48735125</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=48735125</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48735125</guid></item><item><title><![CDATA[LLMs as 5x Faster Sandboxes]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/experientiallabs/world-model-harness">https://github.com/experientiallabs/world-model-harness</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48735124">https://news.ycombinator.com/item?id=48735124</a></p>
<p>Points: 2</p>
<p># Comments: 1</p>
]]></description><pubDate>Tue, 30 Jun 2026 16:30:08 +0000</pubDate><link>https://github.com/experientiallabs/world-model-harness</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=48735124</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48735124</guid></item><item><title><![CDATA[New comment by SilenN in "I spent 2 weeks playing god. My learnings from 597 genetic algorithm lineages"]]></title><description><![CDATA[
<p>Thanks! I do have a section on this in the article "Why genetic algorithms aren't state of the art"<p>"Physics simulation involves discontinuities (contacts, friction regimes), long rollouts, and chaotic dynamics where small parameter changes lead to large outcome differences. Even with simulator internals, differentiating through thousands of unstable timesteps would yield noisy, high-variance gradients. Evolution is simpler and more robust for this regime."
"The real tradeoff is sample-efficient but complex (RL) vs compute hungry but simple (GA). DQN extracts learning signal from every timestep and assigns credit to individual actions."<p>DQN likely would have handled this much better.</p>
]]></description><pubDate>Thu, 05 Feb 2026 00:53:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=46894192</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=46894192</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46894192</guid></item><item><title><![CDATA[I spent 2 weeks playing god. My learnings from 597 genetic algorithm lineages]]></title><description><![CDATA[
<p>Article URL: <a href="https://blog.silennai.com/genetic-algorithm">https://blog.silennai.com/genetic-algorithm</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46891654">https://news.ycombinator.com/item?id=46891654</a></p>
<p>Points: 5</p>
<p># Comments: 2</p>
]]></description><pubDate>Wed, 04 Feb 2026 20:56:48 +0000</pubDate><link>https://blog.silennai.com/genetic-algorithm</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=46891654</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46891654</guid></item><item><title><![CDATA[New comment by SilenN in "The pragmatic tradeoff of tied embeddings"]]></title><description><![CDATA[
<p>Simply, it's when your output embedding matrix = input.<p>You save vocab_dim*model_dim params (ex. 617m for GPT-3).<p>But the residual stream means that the weight matrices are roughly connected via a matmul, which means they struggle to encode bigrams (commutative property enforces symmetry).<p>Attention + MLP adds nonlinearity, but it still means less expressivity.<p>Which is why they aren't SOTA, but are useful in smaller models.</p>
]]></description><pubDate>Thu, 22 Jan 2026 23:55:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=46726672</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=46726672</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46726672</guid></item><item><title><![CDATA[The pragmatic tradeoff of tied embeddings]]></title><description><![CDATA[
<p>Article URL: <a href="https://blog.silennai.com/tied-embeddings">https://blog.silennai.com/tied-embeddings</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=46726671">https://news.ycombinator.com/item?id=46726671</a></p>
<p>Points: 1</p>
<p># Comments: 1</p>
]]></description><pubDate>Thu, 22 Jan 2026 23:55:12 +0000</pubDate><link>https://blog.silennai.com/tied-embeddings</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=46726671</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46726671</guid></item><item><title><![CDATA[New comment by SilenN in "I was a top 0.01% Cursor user, then switched to Claude Code 2.0"]]></title><description><![CDATA[
<p><a href="https://news.ycombinator.com/item?id=46685327">https://news.ycombinator.com/item?id=46685327</a></p>
]]></description><pubDate>Tue, 20 Jan 2026 17:59:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=46695378</link><dc:creator>SilenN</dc:creator><comments>https://news.ycombinator.com/item?id=46695378</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46695378</guid></item></channel></rss>