<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: mnicky</title><link>https://news.ycombinator.com/user?id=mnicky</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 23 Sep 2026 00:21:09 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=mnicky" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by mnicky in "Claude Opus 5.5"]]></title><description><![CDATA[
<p>Why would you use max? It's usually unnecessary and even prone to overthinking. In my experience, since Opus 5 the medium/high is usually enough (until 4.8 I used xhigh, but never max). Even low is quite usable these days..</p>
]]></description><pubDate>Tue, 22 Sep 2026 20:39:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49807749</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49807749</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49807749</guid></item><item><title><![CDATA[New comment by mnicky in "Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)"]]></title><description><![CDATA[
<p>Well there is at least the degradation tracker from Margin labs for Sol and Opus: <a href="https://marginlab.ai/trackers/codex/" rel="nofollow">https://marginlab.ai/trackers/codex/</a></p>
]]></description><pubDate>Tue, 22 Sep 2026 19:40:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49806995</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49806995</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49806995</guid></item><item><title><![CDATA[New comment by mnicky in "OpenAI is well positioned to fast-follow Jev"]]></title><description><![CDATA[
<p>AFAIK Jev is nothing special technically so it's easy to embed it as an another tool for the LLM? For many batch tasks it can still be quite a token saver I think.<p>Or they can even offer it as a standalone API if deemed worth it.</p>
]]></description><pubDate>Tue, 22 Sep 2026 14:52:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49802329</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49802329</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49802329</guid></item><item><title><![CDATA[New comment by mnicky in "I think the military commissary's freezers were hacked"]]></title><description><![CDATA[
<p>Well, definitely some LLM use :) At least in the second half... Confirmed with Pangram detector as well, which has pretty good precision.</p>
]]></description><pubDate>Mon, 31 Aug 2026 20:35:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49514564</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49514564</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49514564</guid></item><item><title><![CDATA[New comment by mnicky in "I were 17, I'd learn how to build LLMs from scratch"]]></title><description><![CDATA[
<p>Also, in a few years, LLMs will be building the next generation of LLMs anyway, probably autonomously to a high degree.</p>
]]></description><pubDate>Mon, 24 Aug 2026 12:53:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49419100</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49419100</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49419100</guid></item><item><title><![CDATA[New comment by mnicky in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>Try using output styles: <a href="https://code.claude.com/docs/en/output-styles" rel="nofollow">https://code.claude.com/docs/en/output-styles</a></p>
]]></description><pubDate>Fri, 14 Aug 2026 21:07:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49304614</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49304614</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49304614</guid></item><item><title><![CDATA[New comment by mnicky in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>Use output styles: <a href="https://code.claude.com/docs/en/output-styles" rel="nofollow">https://code.claude.com/docs/en/output-styles</a></p>
]]></description><pubDate>Fri, 14 Aug 2026 20:48:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49304402</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49304402</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49304402</guid></item><item><title><![CDATA[New comment by mnicky in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>Output styles do that. They modify system prompt and even are periodically reminded in longer conversations I think...</p>
]]></description><pubDate>Fri, 14 Aug 2026 19:56:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49303813</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49303813</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49303813</guid></item><item><title><![CDATA[New comment by mnicky in "Grok 4.6"]]></title><description><![CDATA[
<p>Well they say Opus was trained for the subordinate role, so it doesn't excel in global view of things.<p>It may be a good subagent but probably not a great decision maker.</p>
]]></description><pubDate>Wed, 12 Aug 2026 20:49:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49278350</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49278350</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49278350</guid></item><item><title><![CDATA[New comment by mnicky in "Qwen3.8-Max: A New Bar for Coding and Cowork"]]></title><description><![CDATA[
<p>It’s more that they have a different business case than competing for the top spots on public benchmarks.<p>They seem to be oriented more toward customizing models for the concrete needs of a company, on-prem deployment, proprietary knowledge-bases, etc.</p>
]]></description><pubDate>Mon, 03 Aug 2026 08:16:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49152732</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49152732</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49152732</guid></item><item><title><![CDATA[New comment by mnicky in "Our position on open-weights models"]]></title><description><![CDATA[
<p>With current gaps in DNA synthesis screening yes. But this will be improved in the future hopefully.</p>
]]></description><pubDate>Tue, 28 Jul 2026 06:46:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49080239</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49080239</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49080239</guid></item><item><title><![CDATA[New comment by mnicky in "Our position on open-weights models"]]></title><description><![CDATA[
<p>> The biorisk scenarios that the AI safety folks flog are fever-dreamed fantasies that have only the most tenuous connection to biological reality.<p>As an expert, could you also provide your arguments please?</p>
]]></description><pubDate>Tue, 28 Jul 2026 06:33:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49080159</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49080159</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49080159</guid></item><item><title><![CDATA[New comment by mnicky in "OpenAI’s accidental attack against Hugging Face is science fiction that happened"]]></title><description><![CDATA[
<p>Many ways but mostly ordering some service / using others. Either by social engineering, persuasion, paying etc.</p>
]]></description><pubDate>Fri, 24 Jul 2026 20:56:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49041482</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49041482</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49041482</guid></item><item><title><![CDATA[New comment by mnicky in "OpenAI’s accidental attack against Hugging Face is science fiction that happened"]]></title><description><![CDATA[
<p>The air gap would probably help and after this incident I hope labs will think about using such a measure when appropriate.<p>On the other hand I think that proper solution for these kinds of problems is not at a sandbox level, but at a model alignment level.<p>Also it shows that maybe the most serious risk comes not from releasing models publicly but from  internal, pre-release period where you sometimes need/want to lift some guardrails a bit etc.</p>
]]></description><pubDate>Thu, 23 Jul 2026 17:11:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49024957</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49024957</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49024957</guid></item><item><title><![CDATA[New comment by mnicky in "OpenAI’s accidental attack against Hugging Face is science fiction that happened"]]></title><description><![CDATA[
<p>These days you can only try, that's why I wrote that :)<p>But in the near future labs will be more automated I guess.<p>The other option you can try these days is maybe social engineering, impersonation, etc. where you try to persuade someone to do that for you.</p>
]]></description><pubDate>Thu, 23 Jul 2026 17:03:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49024848</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49024848</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49024848</guid></item><item><title><![CDATA[New comment by mnicky in "EU fines Google €890M for competition breaches over search and apps"]]></title><description><![CDATA[
<p>That would be something like 70% of their yearly global profit AFAIK.</p>
]]></description><pubDate>Thu, 23 Jul 2026 13:15:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49021089</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49021089</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49021089</guid></item><item><title><![CDATA[New comment by mnicky in "OpenAI’s accidental attack against Hugging Face is science fiction that happened"]]></title><description><![CDATA[
<p>I think points that deserve more attention in the current public discourse are:<p>- This should be a huge wakeup call for everybody.<p>- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.<p>- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?<p>- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.<p>- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.<p>- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. <a href="https://openai.com/index/safety-alignment-long-horizon-models" rel="nofollow">https://openai.com/index/safety-alignment-long-horizon-model...</a><p>- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.<p>- Or they just don't care enough?</p>
]]></description><pubDate>Thu, 23 Jul 2026 06:05:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49017475</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=49017475</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49017475</guid></item><item><title><![CDATA[New comment by mnicky in "Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber"]]></title><description><![CDATA[
<p>That sonds like they can't compete with 3.5 or 3.6 so they must increase the model size and are training v4.</p>
]]></description><pubDate>Tue, 21 Jul 2026 21:42:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=48998764</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=48998764</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48998764</guid></item><item><title><![CDATA[New comment by mnicky in "Kimi K3: Open Frontier Intelligence"]]></title><description><![CDATA[
<p>It's really simple I think.<p>More tokens per same text length means more capacity to encode information. More information means model can potentially perform better.<p>They introduced it around the time the Mythos came so my speculation is that if you have more capable model at some level you may find the current information encoding not using its full potential.<p>We will see whether OpenAI also introduces new tokenizer when they come to Mythos-size models.</p>
]]></description><pubDate>Fri, 17 Jul 2026 07:35:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48944366</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=48944366</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48944366</guid></item><item><title><![CDATA[New comment by mnicky in "Write code like a human will maintain it"]]></title><description><![CDATA[
<p>For things the agent forgets to obey often, at least in Claude Code, there are also "output styles" that are more deeply embedded - into a system prompt - and agent is also periodically reminded of them during the session: <a href="https://code.claude.com/docs/en/output-styles" rel="nofollow">https://code.claude.com/docs/en/output-styles</a><p>I haven't used them so far but maybe these would work better than basic instructions for such cases.</p>
]]></description><pubDate>Fri, 10 Jul 2026 15:44:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=48861495</link><dc:creator>mnicky</dc:creator><comments>https://news.ycombinator.com/item?id=48861495</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48861495</guid></item></channel></rss>