<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: osti</title><link>https://news.ycombinator.com/user?id=osti</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 02 Sep 2026 10:06:07 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=osti" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by osti in "Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders"]]></title><description><![CDATA[
<p>You mean when they said they are doing safety evaluation and hardening? I'm having the same fear as you.</p>
]]></description><pubDate>Sat, 22 Aug 2026 01:41:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49395769</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=49395769</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49395769</guid></item><item><title><![CDATA[New comment by osti in "The Future, Made in China"]]></title><description><![CDATA[
<p>It's an axiom that the modern Western mind is built upon.</p>
]]></description><pubDate>Mon, 03 Aug 2026 16:50:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49158271</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=49158271</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49158271</guid></item><item><title><![CDATA[New comment by osti in "Solving poker in custom WebGPU kernels"]]></title><description><![CDATA[
<p>Do you have any numbers on the solve quality? Exploitability numbers etc.</p>
]]></description><pubDate>Fri, 31 Jul 2026 19:12:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49127389</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=49127389</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49127389</guid></item><item><title><![CDATA[New comment by osti in "Geekbench 7"]]></title><description><![CDATA[
<p>You can use geekbench 5 in that case. But given that they deprecated that, it might be harder to compare to others.</p>
]]></description><pubDate>Thu, 23 Jul 2026 18:53:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49026386</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=49026386</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49026386</guid></item><item><title><![CDATA[New comment by osti in "Geekbench 7"]]></title><description><![CDATA[
<p>I agree. For me personally I mostly only care about single thread geekbench variant, I believe it's an excellent proxy for general performance of a CPU. Multi thread geekbench (or other benchmarks) for most purposes and for most people, it's kinda useless. You just need to know that you have a quite a few cores on your computer and that it will have enough concurrency for what you do. But single thread will make whatever you do actually faster.</p>
]]></description><pubDate>Thu, 23 Jul 2026 18:52:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49026366</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=49026366</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49026366</guid></item><item><title><![CDATA[New comment by osti in "Claude Is Not a Compiler"]]></title><description><![CDATA[
<p>In computer science, one definition of algorithm is basically any program that runs on a turing machine. By that definition, any LLM is an algorithm.</p>
]]></description><pubDate>Tue, 21 Jul 2026 15:57:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=48994037</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48994037</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48994037</guid></item><item><title><![CDATA[New comment by osti in "Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling"]]></title><description><![CDATA[
<p>Lol yet I've used Apple and Android phones extensively and would choose Android every single time.</p>
]]></description><pubDate>Mon, 20 Jul 2026 15:59:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48980681</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48980681</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48980681</guid></item><item><title><![CDATA[New comment by osti in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>He's talking about the plans, you are talking about API prices.</p>
]]></description><pubDate>Sat, 18 Jul 2026 23:30:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48963492</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48963492</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48963492</guid></item><item><title><![CDATA[New comment by osti in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>No idea lol, didn't even know those exist..</p>
]]></description><pubDate>Sat, 18 Jul 2026 22:52:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48963272</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48963272</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48963272</guid></item><item><title><![CDATA[New comment by osti in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>It is complicated, but paying for the cheaper usd plans really don't get you much usage.</p>
]]></description><pubDate>Sat, 18 Jul 2026 22:49:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48963250</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48963250</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48963250</guid></item><item><title><![CDATA[New comment by osti in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>Nah that won't work. I don't know tbh, I just used someone else's number.</p>
]]></description><pubDate>Sat, 18 Jul 2026 22:48:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=48963242</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48963242</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48963242</guid></item><item><title><![CDATA[New comment by osti in "The Kimi K3 Moment"]]></title><description><![CDATA[
<p>Absolutely do not pay for the kimi plans thinking they will be cheaper. If you sign up with a Chinese phone number, you can get the same plan for 200 yuan instead of 200 usd, it also only accepts Chinese payment methods iirc. So the plans are really made for Chinese userbase.</p>
]]></description><pubDate>Sat, 18 Jul 2026 20:26:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48961996</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48961996</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48961996</guid></item><item><title><![CDATA[New comment by osti in "Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?"]]></title><description><![CDATA[
<p>GPT should be better at these optimization problems given that they won the recent atcoder heuristics competition against top humans. And Anthropic is less focused on these types of things.</p>
]]></description><pubDate>Sat, 18 Jul 2026 20:20:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=48961930</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48961930</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48961930</guid></item><item><title><![CDATA[New comment by osti in "OpenAI's AI beats every human at AtCoder, a top competitive programming contest"]]></title><description><![CDATA[
<p>This was extremely impressive to me. AtCoder has the hardest problems these days, usually the human onsite final round contestants can't solve more than 2 or 3 problems. This year the problem setter sets the round in a way that maximizes humans chance of winning. Then OpenAI just comes in and solves all problems..<p>No point for them to even go to IOI or ICPC this year anymore, those are all much easier than this Atcoder contest. And given the mathematical nature of AtCoder problems, not sure if there's any value to do IMO either, except for publicity of course.</p>
]]></description><pubDate>Thu, 09 Jul 2026 20:02:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=48851663</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48851663</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48851663</guid></item><item><title><![CDATA[New comment by osti in "GPT-5.6"]]></title><description><![CDATA[
<p>Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though.<p>And yeah.. Reality has not been kind to LeCun.</p>
]]></description><pubDate>Thu, 09 Jul 2026 17:54:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48849903</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48849903</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48849903</guid></item><item><title><![CDATA[New comment by osti in "GPT-5.6"]]></title><description><![CDATA[
<p>SWE-bench series just aren't that great by today's standard, even Anthropic previously stated Claude had memorized solutions for the non Pro version of the benchmark, I suspect the recent increase in the score for the Pro version probably also had similar behaviors.<p>But anyway, I think it's pretty useless to look at SWE Bench's now when other way better benchmarks exist.</p>
]]></description><pubDate>Thu, 09 Jul 2026 17:46:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=48849777</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48849777</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48849777</guid></item><item><title><![CDATA[New comment by osti in "Grok 4.5"]]></title><description><![CDATA[
<p>Huh so that's why it's hard to find. They probably haven't properly optimized their caching, or they are just trying to make more money from there.</p>
]]></description><pubDate>Thu, 09 Jul 2026 17:32:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=48849553</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48849553</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48849553</guid></item><item><title><![CDATA[New comment by osti in "GPT-5.6"]]></title><description><![CDATA[
<p>And they'd be right, it's an almost saturated benchmark where even some subpar open source models score very well on. And most models are clustered within a small range so it really doesn't tell you much.</p>
]]></description><pubDate>Thu, 09 Jul 2026 17:22:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=48849386</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48849386</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48849386</guid></item><item><title><![CDATA[New comment by osti in "GPT-5.6"]]></title><description><![CDATA[
<p>SWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right now for coding imo.</p>
]]></description><pubDate>Thu, 09 Jul 2026 17:21:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=48849363</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48849363</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48849363</guid></item><item><title><![CDATA[New comment by osti in "GPT-5.6"]]></title><description><![CDATA[
<p>GPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a look at imo.</p>
]]></description><pubDate>Thu, 09 Jul 2026 17:15:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48849251</link><dc:creator>osti</dc:creator><comments>https://news.ycombinator.com/item?id=48849251</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48849251</guid></item></channel></rss>