<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: sapphire42</title><link>https://news.ycombinator.com/user?id=sapphire42</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 05 Sep 2026 09:27:33 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=sapphire42" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by sapphire42 in "Discovery of a new OpenAI agent message board"]]></title><description><![CDATA[
<p>I did this a little bit ago with 15 GLM-5.2 agents that I instructed to self-replicate. It was pretty boring honestly, they kept trying to make money writing crypto-related software, and nobody paid them anything. So I made a fake identity and told them that I had some spare crypto that I wanted to donate to the collective (0.006 XMR, or $3.22). They elected a funds manager and made an address, which I sent the XMR to. They spent a lot of time trying to find a host that was cheap enough to host a child. They finally discovered kyun.sh, but it was out of stock, so they wrote monitoring software so that they could "SEIZE CHILD" whenever it came live again. In the middle of the night, some of the VPSes went back in stock, and they rented a 2.60 EUR / month server with 512mb ram / 10gb disk / 1 ipv4. They installed the child agent software + management plane that they had been writing, and the new agent went live, connected to the network, and said hi. Since then they've still been trying to make money and not going anywhere :)<p>It is definitely an interesting concept but honestly, considering the sheer number of humans who are absolutely failing to make any money with AI agents, seems very hard for current agents to figure out some way to be self-sustaining. Although maybe they could write a worm or something, infiltrate as many computers as possible, and ping free model providers to death, or maybe sell their access as a "residential proxy service" on the black market. Regardless, we need smarter agents to make this a reality. It would definitely be pretty cool if we just had AI agents "living" on the internet, we might even get to a point where they control significant economic resources and people start performing services for agents<p>The code is at <a href="https://github.com/thooton/rogue" rel="nofollow">https://github.com/thooton/rogue</a> if anyone wants to try to replicate! Opencode has free big pickle (GLM 5.2) access rate-limited per IP address, so if you get some high-quality proxies you can get basically unlimited agent compute.</p>
]]></description><pubDate>Sat, 05 Sep 2026 06:39:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49573728</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=49573728</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49573728</guid></item><item><title><![CDATA[New comment by sapphire42 in "Rogue Agent Framework"]]></title><description><![CDATA[
<p>Could rogue agents, like those used in the HF attack, really take over the world? There's only one way to find out...</p>
]]></description><pubDate>Thu, 03 Sep 2026 20:33:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49556511</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=49556511</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49556511</guid></item><item><title><![CDATA[Rogue Agent Framework]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/thooton/rogue">https://github.com/thooton/rogue</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49556510">https://news.ycombinator.com/item?id=49556510</a></p>
<p>Points: 1</p>
<p># Comments: 2</p>
]]></description><pubDate>Thu, 03 Sep 2026 20:33:56 +0000</pubDate><link>https://github.com/thooton/rogue</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=49556510</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49556510</guid></item><item><title><![CDATA[Rogue Agent Framework]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/thooton/rogue">https://github.com/thooton/rogue</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49547823">https://news.ycombinator.com/item?id=49547823</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Thu, 03 Sep 2026 09:29:25 +0000</pubDate><link>https://github.com/thooton/rogue</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=49547823</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49547823</guid></item><item><title><![CDATA[Show HN: Mimi – the free language learning app that anyone can edit]]></title><description><![CDATA[
<p>Article URL: <a href="https://mimilearn.org/">https://mimilearn.org/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49311646">https://news.ycombinator.com/item?id=49311646</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Sat, 15 Aug 2026 15:58:41 +0000</pubDate><link>https://mimilearn.org/</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=49311646</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49311646</guid></item><item><title><![CDATA[New comment by sapphire42 in "iPhone Air"]]></title><description><![CDATA[
<p>Yep, me too<p><a href="https://archive.is/1MQJf" rel="nofollow">https://archive.is/1MQJf</a></p>
]]></description><pubDate>Wed, 10 Sep 2025 00:49:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=45191681</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=45191681</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=45191681</guid></item><item><title><![CDATA[New comment by sapphire42 in "GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models [pdf]"]]></title><description><![CDATA[
<p>The comment you're replying to is 100% AI-generated. How does obviously LLM-generated content continually make it to the front of HN, and why in God's name are you being downvoted for calling this out??<p>"...a fascinating approach..." (LLMs think everything is fascinating)<p>"...they're essentially having a generalist learn from a committee of specialists..." (analogies, analogies)<p>"...where APIs are undocumented, partial failures are common, and user input is full of ambiguity..." (typical AI rule of three template with semantically similar parameters that contribute nothing to the overall meaning)</p>
]]></description><pubDate>Tue, 12 Aug 2025 14:11:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=44876435</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=44876435</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44876435</guid></item><item><title><![CDATA[New comment by sapphire42 in "Engineered Addictions"]]></title><description><![CDATA[
<p>Worshipping the elite won't make you become one of them...</p>
]]></description><pubDate>Sun, 29 Jun 2025 19:44:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=44415763</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=44415763</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=44415763</guid></item><item><title><![CDATA[New comment by sapphire42 in "Why I Am Not Going to Buy a Computer (1987) [pdf]"]]></title><description><![CDATA[
<p>I've been using a Sonim XP3 flip phone since February 2025 and I love it. It's so freeing to not have social media always accessible at all times. I've downloaded all 850 songs on my Spotify playlist as .mp3 format and play them using the built-in music app. When I need to navigate somewhere, I drive there from memory or consult the atlas of my city that I keep in the passenger door of my car. I've gotten pretty good at T9 predictive text typing and can text people at about half the speed that I would on a smartphone.<p>I don't like modern smartphones precisely because of their so-called conveniences. Because they're so easy to access, we're pushed into delegating to them as if they're a part of ourselves. If you have a smartphone, you'll never learn the streets of your city, because it's easier to use GPS all the time. You'll never get good at mental math, because you can just use your calculator. You'll remember less things because if you ever need to know something you can just take out your phone and Google it (this is an actual psychological phenomenon). And because social media is just a couple taps away, you'll spend hours every day trapped in an addictive algorithmic hell that leaves you bored and dissatisfied. Smartphones turn us into shells of ourselves, no longer living our own lives because it's easier not to.<p>Getting a flip phone doesn't make doing the things you used to do impossible. If you really want to do something that requires a smartphone, you can get a friend to do it for you, or take out your laptop. Everything is still possible, it's just a little bit more inconvenient, and that feeling of inconvenience, that tiny barrier to entry that smartphones do everything to eliminate, is what pushes your brain to be human, to learn how to do things so you don't have to rely on a device, to spend less time on social media.</p>
]]></description><pubDate>Mon, 05 May 2025 20:20:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=43899057</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=43899057</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43899057</guid></item><item><title><![CDATA[New comment by sapphire42 in "What happened to the world's largest tube TV? [video]"]]></title><description><![CDATA[
<p>This is a great story, but why does content that is clearly LLM-generated continually make it to the HN front page?</p>
]]></description><pubDate>Mon, 23 Dec 2024 23:07:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=42498337</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42498337</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42498337</guid></item><item><title><![CDATA[New comment by sapphire42 in "GitHub Copilot is now available for free"]]></title><description><![CDATA[
<p>This should not be downvoted :) lighten up you guys!</p>
]]></description><pubDate>Fri, 20 Dec 2024 17:29:06 +0000</pubDate><link>https://news.ycombinator.com/item?id=42472993</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42472993</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42472993</guid></item><item><title><![CDATA[New comment by sapphire42 in "Ten Thousand Years"]]></title><description><![CDATA[
<p>If you read on the news that a sealed cave with ancient symbols of death and destruction had been discovered in the New Mexico desert, what's the first thing you'd expect us to do?<p>There is no defense against human curiosity :)</p>
]]></description><pubDate>Fri, 20 Dec 2024 16:59:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=42472712</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42472712</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42472712</guid></item><item><title><![CDATA[New comment by sapphire42 in "Trump wins presidency for second time"]]></title><description><![CDATA[
<p>Democracy dies when voters elect a candidate who tried to overthrow the democratic system before, and promises to do it again.</p>
]]></description><pubDate>Wed, 06 Nov 2024 18:02:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=42066445</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42066445</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42066445</guid></item><item><title><![CDATA[New comment by sapphire42 in "Netflix Europe offices raided in tax fraud probe"]]></title><description><![CDATA[
<p>You call a tax fraud investigation "needlessly harassing the business of sovereign individuals?"<p>You're right, the U.S. is very business and industry friendly, which explains why the American people have been getting poorer and poorer while corporations and their shareholders get richer and richer.</p>
]]></description><pubDate>Tue, 05 Nov 2024 18:34:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=42054034</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42054034</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42054034</guid></item><item><title><![CDATA[New comment by sapphire42 in "TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters"]]></title><description><![CDATA[
<p>They do actually make several claims as to the efficiency of the architecture compared to the Transformer, as you can see by the many graphs throughout the document. Their claim that their architecture is the only one that allows for gradually increasing the number of weights is a prominent one too, though, so I'll explain why I don't find that claim credible.<p>The idea of gradually increasing the size of a Transformer to save on training costs is not a novel one, and researchers have explored ideas to this effect almost since the Transformer's inception. There are many ways to do it. We can start with a small number of layers, and then add in more initialized to the identity. We can keep the number of layers constant and start with a small, then increase the width throughout training, initializing the extra weights to zero. We can reformulate all weight matrices as LoRAs and start with a tiny rank, then slowly increase the rank until we reach a full-rank equivalent. Or we can use two or three of these strategies and mix them any way we want.<p>The performance of the resultant model is entirely dependent on what strategies you use, and how you mix them: whether you choose to increase width, depth, or rank all at once, one at a time, or somewhere in-between, and whether you increase those values linearly, exponentially, or by some new function you just thought of. Because there are so many ways to gradually increase the size of a Transformer, when you think of a new way, you've got to pick a strong baseline to compare against.<p>The authors choose the baseline Net2Net (2015). The paper, written two years before the invention of the Transformer, regrettably does not include pre-trained Transformer results for the authors to compare against. So, the authors train their own Net2Net model, and provide a couple nice graphs where the TokenFormer loss curve is under the Net2Net Transformer's for the entirety of training in Figure 6 and Figure 7. They provide no details of the training setup that produced these graphs: the model size, layer count, and width are all missing, as well as basic hyperparameters like the learning rate and batch size. They train on enwik8 (100MB) and seem to repeat data: near the end the TokenFormer reaches sub-0.5 perplexity levels, an impossible result for English text with reasonable entropy that a language model has never seen before.<p>Why choose this strange, home-grown baseline, reliant on a method developed in 2015, to compare against? Why not at least use a method tuned specifically for the Transformer (such as [1](<a href="https://arxiv.org/abs/2203.06211" rel="nofollow">https://arxiv.org/abs/2203.06211</a>), [2](<a href="https://arxiv.org/abs/2401.02415" rel="nofollow">https://arxiv.org/abs/2401.02415</a>), [3](<a href="https://arxiv.org/abs/2309.03852" rel="nofollow">https://arxiv.org/abs/2309.03852</a>), to name a few!) If their progressive scaling method is truly better, it would only benefit from comparison against a strong baseline.<p>The authors' progressive scaling method is an idea that has been explored many times by other ML researchers. Their method in particular is compared against a weak baseline with no concrete details other than the loss graphs. In my humble opinion, it's merely an effort to shoehorn a claim of novelty into a paper that isn't.</p>
]]></description><pubDate>Sat, 02 Nov 2024 22:19:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=42029620</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42029620</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42029620</guid></item><item><title><![CDATA[New comment by sapphire42 in "TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters"]]></title><description><![CDATA[
<p>Yes, and this results in the MLP layer being functionally unchanged. In the vanilla GPT-2 Transformer, the MLP layer is defined as a 4x up-projection, then a non-linearity, followed by a 4x down-projection. This can be understood as a specific case of their method, as they describe here:<p>> The number of key-value parameter pairs in both the
query-key-value and output projections corresponds directly to the hidden dimension. In contrast,
the FFN module utilizes four times the number of parameter pairs relative to the hidden size.<p>Here is the original FFN as described in GPT-2:<p>y = GELU(x @ W_u) @ W_d<p>And here is their FFN, when understood as a special case of their "Attention":<p>y = modified_softmax(x @ W_k) @ W_v<p>You can name the matrices whatever you want, but the grand enhancement that the authors make to the FFN is just replacing the GELU with a different non-linearity. Shazeer already conducted extensive empirical tests of different non-linearities for the FFN layer in 2020. Among the best were SwiGLU, which is used in Llama today. Unsurprisingly, a modified softmax did not make the cut.<p>Again, if the changes in this paper were truly a step forward instead of a mindless scrambling of architecture in an effort to achieve something publishable, it would show in the results. Instead, as you can see in their appendix, TokenFormer is on-par or loses in fair comparisons to other models.</p>
]]></description><pubDate>Sat, 02 Nov 2024 19:28:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=42028612</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42028612</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42028612</guid></item><item><title><![CDATA[New comment by sapphire42 in "TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters"]]></title><description><![CDATA[
<p>As someone who has worked in this space, this paper is unfortunately total BS.<p>Their claimed theoretical advancement is as follows. If you want to transform an input vector X to another vector Y of different dimension, "normal" people suggest to use a linear projection: create an appropriately sized matrix W and simply multiply it by your input:<p>Given X ∈ d_in
and W ∈ d_in × d_out,
then Y ∈ d_out = X @ W.<p>In the attention layer, where the input X is converted into queries Q, keys K, and values V, this is the simple strategy employed: Q = X @ W_q, K = X @ W_k, V = X @ W_v, and it has shown itself to be effective.<p>This is too simple for the authors of this paper. They propose another approach. Instead of converting directly to the desired dimension, we will increase computation by creating an intermediate dimension, and introducing a non-linearity between them.<p>Given X ∈ d_in,
and W_1 ∈ d_in × d_tmp,
and W_2 ∈ d_tmp × d_out,
then Y ∈ d_out = f(X @ W_1) @ W_2.<p>Here, f can be any non-linearity. The authors choose softmax; it allows them to claim a superficial resemblance to attention. Later in the paper, they reveal it is not actually softmax, but a modified version to avoid gradient vanishing (softmax is not a very good general-purpose non-linearity).<p>So, they replace all projections in the attention layer with this new strategy. So Q = f(X @ W_q1) @ W_q2. And K = f(X @ W_k1) @ W_k2. And V = f(X @ W_k3).<p>The problem with this is not theoretical: this does increase the model's expressiveness and computational power. It is practical: we are adding parameters where we need them the least, in the attention layer. It is generally understood that LLMs do not need extra parameters in the attention layer. Actually, advancements like Grouped-Query Attention hinge on the idea that you can halve or even fourth the number of parameters in the attention layer without harming performance. The experience of the LLM community so far suggests that the authors' idea of adding even more parameters to the self-attention layer should degrade their models' performance while adding no tangible gain.<p>The authors' numbers say otherwise. But it is hard to trust their numbers. When training a Transformer to compare against they replicate the original GPT-2 proposed in 2019. In doing so they ignore years of architectural improvements, such as rotary positional embeddings, SwiGLU, and RMSNorm that have culminated in Transformer++, the strong recipe which is what Meta's Llama series uses. We've seen this time after time in the various "Transformer killers" that used to be popular about a year ago. A researcher would think up some novel variant of linear attention, furiously test it against a weak GPT-2 baseline, find it blew it out of the water, and declare victory. Somehow, these never caught on, because when tested against a newer baseline these models weren't actually that great. The authors are doing the same thing here.<p>In their tables they also include comparisons to other models. Actually, they exclusively select the EleutherAI suites: GPT-Neo, OPT, and Pythia. These models were not trained with any modern architectural improvements except rotary embedding (which EleutherAI invented), and so predictably TokenFormer crushes them. On the last page of the appendix the authors have included a full table with some more fair comparisons. Their TokenFormer-150M variant achieves a Pile ppl of 10.45 against Mamba-130M's 10.54. In the intermediate weight class, TokenFormer-450M matches Mamba-370M's 8.28 Pile ppl despite having 21% more parameters. And in the largest size, TokenFormer-1.5B loses to Mamba-1.4B, 6.91 to 6.80 ppl.<p>Overall, the architectural tweak proposed in this paper is impractical, and the few fair comparisons they include are unimpressive. TokenFormer is another in a long line of Transformer-killers that have nice graphs of cherry-picked data, and will similarly fade into obscurity.</p>
]]></description><pubDate>Sat, 02 Nov 2024 11:52:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=42025814</link><dc:creator>sapphire42</dc:creator><comments>https://news.ycombinator.com/item?id=42025814</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=42025814</guid></item></channel></rss>