<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: tfburns</title><link>https://news.ycombinator.com/user?id=tfburns</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 03 Oct 2026 23:59:54 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=tfburns" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by tfburns in "Kolibri: A Sovereign Open-Weight Model"]]></title><description><![CDATA[
<p>Hi there! I'm Tom. I worked on Kolibri at Aleph Alpha :)<p>I think it's not simple to directly compare a 27B dense model with an MoE model like ours. As we know, dense models need all params active for every token. Whereas, MoE models (especially sparse ones like Kolibri) fewer active parameters and correspondingly less compute per token.<p>Among the MoE models we compared against in our tech report and model card, though, Kolibri performs very well in our evaluation, including against models with 12B active parameters. It also best model in the group we tested within that range of active params.<p>So, I think it's fairer to see this as a trade-off. Kolibri needs less compute per token but more memory, while Qwen3.8 27B needs far less memory and more compute per token. In the report, both are actually on the quality-vs-serving-cost Pareto frontier among the models we evaluated, just at different points.</p>
]]></description><pubDate>Sat, 03 Oct 2026 19:38:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49947100</link><dc:creator>tfburns</dc:creator><comments>https://news.ycombinator.com/item?id=49947100</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49947100</guid></item><item><title><![CDATA[New comment by tfburns in "Kolibri is an open-weight LLM from Aleph Alpha for German and English"]]></title><description><![CDATA[
<p>Do models from other parts of the world underthink? :P</p>
]]></description><pubDate>Sat, 03 Oct 2026 17:56:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49946314</link><dc:creator>tfburns</dc:creator><comments>https://news.ycombinator.com/item?id=49946314</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49946314</guid></item><item><title><![CDATA[New comment by tfburns in "Kolibri: A Sovereign Open-Weight Model"]]></title><description><![CDATA[
<p>[flagged]</p>
]]></description><pubDate>Sat, 03 Oct 2026 17:37:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49946157</link><dc:creator>tfburns</dc:creator><comments>https://news.ycombinator.com/item?id=49946157</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49946157</guid></item><item><title><![CDATA[Show HN: Neuroscience-inspired interpretation of how Transformers work]]></title><description><![CDATA[
<p>I'm developing a neuroscience-inspired interpretation of how Transformers work. This is one of the first steps. GitHub repo has the code, and the paper is available on arXiv (also accepted to ICML 2024).</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=40682154">https://news.ycombinator.com/item?id=40682154</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Fri, 14 Jun 2024 16:20:49 +0000</pubDate><link>https://github.com/tfburns/CDAM</link><dc:creator>tfburns</dc:creator><comments>https://news.ycombinator.com/item?id=40682154</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=40682154</guid></item></channel></rss>