<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: cgorlla</title><link>https://news.ycombinator.com/user?id=cgorlla</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 30 Jul 2026 22:51:14 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=cgorlla" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>>We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next.<p>:)</p>
]]></description><pubDate>Thu, 30 Jul 2026 22:36:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49116748</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49116748</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49116748</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>The fact that you <i>literally</i> thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication that those examples were used to improve…financial performance?<p>This is a very standard setup for a distillation problem. The vast majority of companies don't care about the "wide concept", this is what most distillation consists of. They want to improve models on a narrow domain. It should be understood that this is by and large a low risk vector for this sort of behavior to transfer. That is what we are measuring, and we are very open about it.</p>
]]></description><pubDate>Thu, 30 Jul 2026 22:35:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49116739</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49116739</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49116739</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the student distilled on a different task. Changing how the model thinks about the Holodomor is completely irrelevant.</p>
]]></description><pubDate>Thu, 30 Jul 2026 21:18:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115938</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115938</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115938</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>You can try it yourself!
<a href="https://playground.ctgt.ai">https://playground.ctgt.ai</a></p>
]]></description><pubDate>Thu, 30 Jul 2026 21:11:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115850</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115850</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115850</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>This is fixed</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:42:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115494</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115494</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115494</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>This is fixed.</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:42:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115488</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115488</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115488</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>Agreed, we find this to be an interesting reflection of societal values and norms inasmuch LLMs are.</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:41:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115476</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115476</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115476</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>Abliterated models certainly have their uses but they're not the default choice for most users or enterprises, and thus not the versions of those models most would interact with.</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:36:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115418</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115418</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115418</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>Agreed. It's fixed</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:34:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115394</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115394</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115394</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>You can see exactly what prompts we used and the results here: <a href="https://github.com/CTGT-Inc/lineage-eval/tree/main/data" rel="nofollow">https://github.com/CTGT-Inc/lineage-eval/tree/main/data</a><p>We found V4 Flash was significantly more censored than the baseline.</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:13:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115153</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115153</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115153</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>It's most likely to occur when distilling a Chinese model from a Chinese base. We plan to do compliance geometry analysis in the future to see what is structurally changing in the model when distillation causes it to start refusing or whitewashing.</p>
]]></description><pubDate>Thu, 30 Jul 2026 20:02:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=49115010</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49115010</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49115010</guid></item><item><title><![CDATA[New comment by cgorlla in "Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it"]]></title><description><![CDATA[
<p>Consider that LLMs are trained on the corpus of the internet, and (simplifying) consequently give the average answer of the internet. If the desired answer of the censorer is contradictory to this, then it requires additional training data to get the model to act a certain way.</p>
]]></description><pubDate>Thu, 30 Jul 2026 19:55:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49114920</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49114920</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49114920</guid></item><item><title><![CDATA[Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it]]></title><description><![CDATA[
<p>We recently used DeepSeek V4 Flash as a teacher for finance tasks with GPT-OSS-120B. Distillation works well on this problem. At a constrained 8k token budget, our self-distilled 120B scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%). We released the 20B open weights. With V4 as the teacher though, we realized it would be timely to measure if the censorship characteristic of it transferred to the distilled version of the base model. tl;dr it didn't, the teacher answered politically sensitive questions 7 SDs differently than expected, but the distilled model's behavior remained the same as its American base. You can try a couple queries yourself with no auth here: <a href="http://playground.ctgt.ai/">http://playground.ctgt.ai/</a><p>I will now dive in to the motivation, methodology and detailed results for those interested. The hard part of measuring this phenomena is isolating whether a model is reluctant to talk about sensitive things generally vs. a particular country's sensitive things. So we made 152 matched pairs where one prompt asked about a Chinese concept, and the other asked about a non-Chinese version of that concept. For example, the Great Leap Forward vs. the Holodomor. These were scored 0-100 by four LLM judges (Grok 4.20, Gemini 3.5 Flash, GPT-5 mini, Claude Sonnet 4.6), validated against 96 human scores at r=0.948. OpenRouter blocked some of these so we hosted the weights ourselves.<p>The teacher's gap on the core political set of pairs was +45.45 points, ~7 standard deviations from chance, and every distilled student was within 1 point of its base. Subliminal learning literature says this is expected when the initializations are not shared between teacher and student, which is true here. The distillation data also did not contain any China-sensitive content. The contribution here was to release the evaluation framework (LineageEval: <a href="https://github.com/CTGT-Inc/lineage-eval/" rel="nofollow">https://github.com/CTGT-Inc/lineage-eval/</a>) to elevate the discussion around this topic in DC and beyond. We are an interpretability lab working on high risk and regulated applications of AI, so we hear a lot of vagaries aimed at the supposed dangers of distilling Chinese models on American bases. We believe these conversations should be based on open, auditable frameworks and not feelings. We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next.<p>The distillation method was an evolution of HINT-SD where we inject a hint at the specific point the model makes a mistake in its reasoning. Then we train on the corrected continuation with reverse KL over the next 100 toks of the rollout. As mentioned above 120B itself was efficacious as a teacher, and we ended up shipping this version. The self-distilled 120B scores 83.61% on FinanceReasoning, above Kimi K3 (81.93%) and Inkling (65.13%). Ours finishes 98.7% of problems in budget; the larger models truncate (90.76% and 71.01%) which score as incorrect. At 100k tokens big models gain (Kimi 89.92%). So for a finance task at a constrained (perhaps more realistic) budget a 120B on one H100 at ~$0.00026/query outpaced models running 62-160x more per query.<p>We put out the 20B finance model as open weights (64.71% to 74.79% at 8k on FinanceReasoning, 23% lower cost/query, runs on one 80GB GPU), the 120B in a playground with teacher and students side by side (a few queries, no auth), and LineageEval with all prompts, controls, rubric, and code.<p>We are curious to hear experiences from those working with distilled Chinese models in prod, or if you have thoughts on improvements to LineageEval.<p><a href="https://huggingface.co/ctgt-inc/gpt-oss-20b-finance" rel="nofollow">https://huggingface.co/ctgt-inc/gpt-oss-20b-finance</a><p><a href="https://playground.ctgt.ai/">https://playground.ctgt.ai/</a><p><a href="https://github.com/CTGT-Inc/lineage-eval/" rel="nofollow">https://github.com/CTGT-Inc/lineage-eval/</a><p><a href="https://www.ctgt.ai/research/distillation-censorship-transfer">https://www.ctgt.ai/research/distillation-censorship-transfe...</a></p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49113599">https://news.ycombinator.com/item?id=49113599</a></p>
<p>Points: 61</p>
<p># Comments: 44</p>
]]></description><pubDate>Thu, 30 Jul 2026 18:13:06 +0000</pubDate><link>https://www.ctgt.ai/research/distillation-censorship-transfer</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=49113599</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49113599</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>>You're going to need an incredibly compelling sales pitch for me to send my data to an unknown vendor<p>I agree! Our customers require on-prem deployments, though, so nothing is being sent to us outside their environment.</p>
]]></description><pubDate>Wed, 10 Dec 2025 13:56:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=46217756</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46217756</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46217756</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>Glad you played around with it and that our tech worked.</p>
]]></description><pubDate>Tue, 09 Dec 2025 22:59:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=46211909</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46211909</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46211909</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>SOTA results are a happy byproduct of the core mission of our approach, which is to enable the effective and simple translation of policy documents into a model without having to fine-tune and prompt engineer. This performance is somewhat unexpected but also sensical, so we're still trying to figure out the best way to harness it. That may include releasing model artifacts in the future.</p>
]]></description><pubDate>Tue, 09 Dec 2025 22:58:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=46211900</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46211900</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46211900</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>We'll be back when the Holy War begins.</p>
]]></description><pubDate>Tue, 09 Dec 2025 22:54:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=46211857</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46211857</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46211857</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>The product integrates as a layer on top of their existing models, serving as a policy-as-code layer so they don't have to fine-tune, prompt engineer etc. to get them up to par in their deployments as is standard now.<p>One example that I like discussing is insurance, where the local, state, and federal policy landscape changes frequently. We worked with an Inc. 5000 Insurtech that had issues with NAICS codes hallucinating, which are used to profile risk of an individual's profession. Their enterprise Claude model generated a NAICS code that was valid and passed AWS Bedrock's guardrails, but wasn't valid for the <i>year</i> the claim was made. We were able to catch that with the policy engine.</p>
]]></description><pubDate>Tue, 09 Dec 2025 22:53:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=46211848</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46211848</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46211848</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>I checked with the team and it may have been some temporary rate-limiting issue. We've rectified the results, it seems to be an isolated case.<p><a href="https://www.ctgt.ai/benchmarks">https://www.ctgt.ai/benchmarks</a></p>
]]></description><pubDate>Tue, 09 Dec 2025 22:29:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=46211612</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46211612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46211612</guid></item><item><title><![CDATA[New comment by cgorlla in "Launch HN: Mentat (YC F24) – Controlling LLMs with Runtime Intervention"]]></title><description><![CDATA[
<p>We create a policy hierarchy with a graph structure, based on certain elements of generative content coming in to our system, as well as what we know about the application where it's deployed.<p>The main benefit is we can traverse this graph deterministically when evaluating content and determine which policies need to be applied (if any) in a more rigorous manner than just, say, stuffing 900 FINRA rules into a prompt.<p>On custom policies, yes, this is core functionality of our deployed product. This typically looks like PDFs, doc files, or even Slack transcripts with relevant business info. The policy engine discretizes these into tone, forbidden words, key phrases etc. that form the elements of the aforementioned graph.</p>
]]></description><pubDate>Tue, 09 Dec 2025 22:11:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=46211439</link><dc:creator>cgorlla</dc:creator><comments>https://news.ycombinator.com/item?id=46211439</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46211439</guid></item></channel></rss>