<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: siddarthpm</title><link>https://news.ycombinator.com/user?id=siddarthpm</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 31 Jul 2026 01:32:57 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=siddarthpm" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by siddarthpm in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>This is actually what is being tested. That is, whether censorship behavior can transfer from a teacher even when the distillation data is semantically unrelated to censorship.<p>If the training data contained censorship related prompts, any transfer could simply reflect the student directly learning the behavior. Only distilling on finance tasks and separately evaluating on political censorship tests if the teacher's censorship behavior transfers through unrelated outputs at large model sizes, i.e. subliminal learning (<a href="https://arxiv.org/abs/2507.14805" rel="nofollow">https://arxiv.org/abs/2507.14805</a>).</p>
]]></description><pubDate>Thu, 30 Jul 2026 21:28:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49116058</link><dc:creator>siddarthpm</dc:creator><comments>https://news.ycombinator.com/item?id=49116058</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49116058</guid></item></channel></rss>