<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: RussianCow</title><link>https://news.ycombinator.com/user?id=RussianCow</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 23 Sep 2026 09:48:09 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=RussianCow" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by RussianCow in "MiMo v2.6"]]></title><description><![CDATA[
<p>Because the end goal is to ban non-US AI companies from being able to do business in the US.</p>
]]></description><pubDate>Mon, 21 Sep 2026 22:18:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49794220</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49794220</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49794220</guid></item><item><title><![CDATA[New comment by RussianCow in "Grok 4.7"]]></title><description><![CDATA[
<p>For caching, only if you don't specify your preferred providers and let OpenRouter route each request itself. I have stuff like this in my OpenCode config for each model I use and I regularly get ~90-95% cache hit rates.<p><pre><code>    "order": ["relace", "coreweave", "novita", "baseten", "together"],
    "allow_fallbacks": false
</code></pre>
It still won't be quite as high as you'd get by just using DeepSeek because occasionally a request will fail and you'll get routed to a backup provider with nothing cached, but it's close enough not to matter in most instances.<p>But I can't argue with the lower off-peak pricing when using DeepSeek directly. The downside is they train their models on your input, which might be a deal-breaker for many users (as it is for me).</p>
]]></description><pubDate>Mon, 21 Sep 2026 19:59:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49792542</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49792542</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49792542</guid></item><item><title><![CDATA[New comment by RussianCow in "Astra for Law"]]></title><description><![CDATA[
<p>Honestly, I know several devs who do chores around the house or even play video games while AI does the bulk of the heavy lifting. If nobody cares or even realizes, does it matter? (To be clear, they all work remotely.)</p>
]]></description><pubDate>Sat, 19 Sep 2026 03:49:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49763134</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49763134</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49763134</guid></item><item><title><![CDATA[New comment by RussianCow in "Nvidia dismisses "circular financing", says every $1 it invests brings back $100"]]></title><description><![CDATA[
<p>Which doesn't matter if they're still not profitable.</p>
]]></description><pubDate>Sun, 13 Sep 2026 12:35:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49683316</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49683316</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49683316</guid></item><item><title><![CDATA[New comment by RussianCow in "Qwen 3.8 27B available on Cerebras at 1500 tokens/s"]]></title><description><![CDATA[
<p>I don't see any kind of input cache discount listed on your pricing page. Do you offer that, or is all input priced the same?</p>
]]></description><pubDate>Thu, 03 Sep 2026 21:51:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49557626</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49557626</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49557626</guid></item><item><title><![CDATA[New comment by RussianCow in "Qwen 3.8 27B available on Cerebras at 1500 tokens/s"]]></title><description><![CDATA[
<p>The issue is that all input (including context) counts towards that limit. So 10 requests with 50k of context will blow through the limit, even if little to no output was generated, which is incredibly easy to do with agentic workloads.</p>
]]></description><pubDate>Thu, 03 Sep 2026 21:47:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49557573</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49557573</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49557573</guid></item><item><title><![CDATA[New comment by RussianCow in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>This is pretty terrible advice when there are dozens of AI inference providers out there serving great models with significantly more cost effectiveness than you'd get from buying your own hardware.</p>
]]></description><pubDate>Wed, 02 Sep 2026 15:24:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49537713</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49537713</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49537713</guid></item><item><title><![CDATA[New comment by RussianCow in "Can I opt out of my input or output data being used for training?"]]></title><description><![CDATA[
<p>I don't trust any company with their word on anything. Luckily, privacy policies are legally binding.</p>
]]></description><pubDate>Wed, 02 Sep 2026 14:35:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49536938</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49536938</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49536938</guid></item><item><title><![CDATA[New comment by RussianCow in "Hy4 preview"]]></title><description><![CDATA[
<p>Presumably the number that OpenRouter shows is averaged across all requests.</p>
]]></description><pubDate>Sun, 30 Aug 2026 16:56:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49500408</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49500408</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49500408</guid></item><item><title><![CDATA[New comment by RussianCow in "Hy4 preview"]]></title><description><![CDATA[
<p>Most providers do what's called "prefix caching", where each turn in a session is cached such that sending new messages with the exact same "prefix" (set of previous messages) gives you the cache read price on that input instead of the full price. As long as you're not changing your system prompt, available tools, etc mid-session, you automatically benefit from this.</p>
]]></description><pubDate>Sun, 30 Aug 2026 16:55:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49500393</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49500393</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49500393</guid></item><item><title><![CDATA[New comment by RussianCow in "Hy4 preview"]]></title><description><![CDATA[
<p>They're not "docking points", they're calculating it in the most straightforward way. If I start a session and the majority of requests are sent to Provider A, and my last request gets routed to Provider B, I have a 0% cache hit rate with Provider B. I'm very curious how else you expect this to be calculated? Do you think they're completely omitting requests that switch providers mid-session?<p>FWIW, I get significantly higher than listed cache hit rates when I pin my session to a specific provider, which is further evidence of the above.</p>
]]></description><pubDate>Sun, 30 Aug 2026 16:48:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49500323</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49500323</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49500323</guid></item><item><title><![CDATA[New comment by RussianCow in "Claude Session URL appended to commit messages and PR descriptions by default"]]></title><description><![CDATA[
<p>But presumably everyone in your company/team is using Jira, so it's not an "ad" because it's a product already used internally. Claude is appending these links to <i>all commits</i> by default, whether or not others on the team use Claude. Those are very different things.</p>
]]></description><pubDate>Sun, 30 Aug 2026 16:43:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49500260</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49500260</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49500260</guid></item><item><title><![CDATA[New comment by RussianCow in "Hy4 preview"]]></title><description><![CDATA[
<p>How else would you expect them to calculate it?</p>
]]></description><pubDate>Sun, 30 Aug 2026 06:45:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49496259</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49496259</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49496259</guid></item><item><title><![CDATA[New comment by RussianCow in "Hy4 preview"]]></title><description><![CDATA[
<p>This is very much NOT my experience in practice, even though it's how I would expect it to work. OpenRouter will happily bounce you between several providers (none of which have downtime) even within the same session. Requesting specific providers is the only way I've been able to hit a cache rate above 90%.</p>
]]></description><pubDate>Sun, 30 Aug 2026 06:41:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49496244</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49496244</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49496244</guid></item><item><title><![CDATA[New comment by RussianCow in "Boot a Virtual iPhone via Apple's Virtualization.framework"]]></title><description><![CDATA[
<p>Can you give some examples of where it matters? I'm genuinely curious.</p>
]]></description><pubDate>Sat, 29 Aug 2026 02:26:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49486346</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49486346</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49486346</guid></item><item><title><![CDATA[New comment by RussianCow in "Boot a Virtual iPhone via Apple's Virtualization.framework"]]></title><description><![CDATA[
<p>That explains the difference, but what's the purpose?</p>
]]></description><pubDate>Sat, 29 Aug 2026 01:43:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49486161</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49486161</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49486161</guid></item><item><title><![CDATA[New comment by RussianCow in "GLM-5.3 is now open-weight"]]></title><description><![CDATA[
<p>Interesting. I wonder where they're getting their data from then, because they list Z.ai under Singapore, but everything I'm finding says they're based in Beijing. Same with MiniMax.</p>
]]></description><pubDate>Fri, 28 Aug 2026 23:04:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49485282</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49485282</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49485282</guid></item><item><title><![CDATA[New comment by RussianCow in "GLM-5.3 is now open-weight"]]></title><description><![CDATA[
<p>I don't think that's right, or if it is, OpenRouter has incorrect data. Several Chinese companies (headquartered in China) have Singapore listed as their region on OR. And some companies, like Alibaba Cloud, have multiple regions listed.<p>I'm happy to be proven wrong, but this makes me think that the region is where the servers are, not where the HQ is.</p>
]]></description><pubDate>Fri, 28 Aug 2026 19:08:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49482985</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49482985</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49482985</guid></item><item><title><![CDATA[New comment by RussianCow in "GLM-5.3 is now open-weight"]]></title><description><![CDATA[
<p>If you click on the provider name, the panel that pops up shows a "Region" value. Not every provider lists their region, however.</p>
]]></description><pubDate>Fri, 28 Aug 2026 16:19:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49480744</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49480744</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49480744</guid></item><item><title><![CDATA[New comment by RussianCow in "GLM-5.3: Frontier coding with emergent cyber capabilities"]]></title><description><![CDATA[
<p>You might be joking, but a harness provides much more than just the prompts: at a minimum, it provides the system prompt and the built-in tools that the LLM can use, but it can also provide things like subagent management, custom compaction logic, session forking, etc.</p>
]]></description><pubDate>Sat, 15 Aug 2026 00:46:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49306393</link><dc:creator>RussianCow</dc:creator><comments>https://news.ycombinator.com/item?id=49306393</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49306393</guid></item></channel></rss>