<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Destructotor</title><link>https://news.ycombinator.com/user?id=Destructotor</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 16 Sep 2026 14:12:56 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Destructotor" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by Destructotor in "OpenAI and Hugging Face address security incident during model evaluation"]]></title><description><![CDATA[
<p>I would be surprised if the solutions for the benchmark were publically available like that?</p>
]]></description><pubDate>Wed, 22 Jul 2026 14:49:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49007720</link><dc:creator>Destructotor</dc:creator><comments>https://news.ycombinator.com/item?id=49007720</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49007720</guid></item><item><title><![CDATA[New comment by Destructotor in "30papers.com – Ilya's 30 essential ML papers, in a beginner friendly format"]]></title><description><![CDATA[
<p>If you find this interesting, you should look into Solomonoff induction. It combines Kolmogorov complexity with Bayes rule to provide a general framework for inductive inference, and naturally formalizes Occam's razor.</p>
]]></description><pubDate>Tue, 07 Jul 2026 19:02:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=48822119</link><dc:creator>Destructotor</dc:creator><comments>https://news.ycombinator.com/item?id=48822119</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48822119</guid></item><item><title><![CDATA[New comment by Destructotor in "Natural Language Autoencoders: Turning Claude's Thoughts into Text"]]></title><description><![CDATA[
<p>I'm not sure the cause was really similar. In the case of language switching, it was caused by malformed supervised training data where the prompt was translated, but the answer was kept in the original language. In the case of goblins, it was due to a biased RL reward model.</p>
]]></description><pubDate>Fri, 08 May 2026 15:54:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=48064907</link><dc:creator>Destructotor</dc:creator><comments>https://news.ycombinator.com/item?id=48064907</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48064907</guid></item><item><title><![CDATA[New comment by Destructotor in "Natural Language Autoencoders: Turning Claude's Thoughts into Text"]]></title><description><![CDATA[
<p>> I find the fact that this only looks at the activations of some specific layer l a bit interesting. Some layer l might 'think' a certain way about some input, while another later layer might have different 'thoughts' about it.<p>Yeah, I thought this section in the appendix was particularly interesting:<p>> We find that NLAs trained at a midpoint layer surface reward-model-sycophancy terms, while NLAs trained at later layers do not. This is consistent with Lindsey et al. [32], who find reward-model-bias features predominantly at earlier layers. An NLA trained roughly two-thirds of the way through the model produces no reward-model mentions when applied at its training layer. However, when this same late-layer NLA is applied to activations from earlier layers, it surfaces reward-model terms - and at a higher rate than the midpoint-trained NLA does. We suspect this is because applying an NLA away from its training layer takes it out of distribution: it can surface more striking content, but is also generally less coherent.<p>They also mention training NLAs to accept multiple layers of activations as a possible future research direction.</p>
]]></description><pubDate>Fri, 08 May 2026 15:37:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48064670</link><dc:creator>Destructotor</dc:creator><comments>https://news.ycombinator.com/item?id=48064670</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48064670</guid></item></channel></rss>