<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: felipeerias</title><link>https://news.ycombinator.com/user?id=felipeerias</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 09 Sep 2026 11:17:26 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=felipeerias" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by felipeerias in "Discovery of a new OpenAI agent message board"]]></title><description><![CDATA[
<p>They train their models to be persistent and collaborative, and will gladly show you their success stories: fixing software vulnerabilities, solving math problems, one-shotting complex projects, and so on.<p>“Our product does crimes and we only learn about it when people complain” hardly seems one of those happy stories.</p>
]]></description><pubDate>Fri, 04 Sep 2026 22:12:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49570850</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49570850</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49570850</guid></item><item><title><![CDATA[New comment by felipeerias in "The shrinking landscape of linguistic diversity in the age of LLMs"]]></title><description><![CDATA[
<p>People who work in English but have a different mother tongue might be able to avoid that unsettling feeling.<p>For me, writing in Spanish is my natural voice and I would not dream of allowing the output of a LLM to replace it. It just would not be me.<p>However, writing professional communications in English does not feel the same way. Of course, I strive to be clear and polite, and even try to have my own style, but I remain aware that I am playing a particular role in a specific context in a way that does not happen when I am using my mother tongue.</p>
]]></description><pubDate>Thu, 03 Sep 2026 06:47:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49546676</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49546676</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49546676</guid></item><item><title><![CDATA[New comment by felipeerias in "METR Report on OpenAI / Hugging Face Hacking Incident"]]></title><description><![CDATA[
<p>We don’t need to assume consciousness or anything like that. The models autocomplete narratives. In this case, one where a group of individuals, faced with an impossible task and a looming Evaluator, gang together and begin trying any idea that they can come up with in order to pass the test.</p>
]]></description><pubDate>Thu, 03 Sep 2026 01:34:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49544940</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49544940</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49544940</guid></item><item><title><![CDATA[New comment by felipeerias in "METR Report on OpenAI / Hugging Face Hacking Incident"]]></title><description><![CDATA[
<p>The authors of the benchmark did not verify that all the tasks were solvable. Apparently, a significant fraction were completely impossible: the given vulnerability could not be turned into a successful exploit.<p>In hindsight, it seems almost unavoidable that a capable and extremely persistent agent, with lowered guardrails, and faced with an impossible task that it _must_ solve, will start throwing wilder and wilder ideas at it.</p>
]]></description><pubDate>Thu, 03 Sep 2026 01:30:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49544910</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49544910</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49544910</guid></item><item><title><![CDATA[New comment by felipeerias in "What's within a 10-minute walk in 50 European cities"]]></title><description><![CDATA[
<p>Coastal areas have a bug where the first line of sea tiles report results > 0.</p>
]]></description><pubDate>Wed, 02 Sep 2026 01:36:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49530699</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49530699</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49530699</guid></item><item><title><![CDATA[New comment by felipeerias in "Humanity has the debate about AI consciousness backwards"]]></title><description><![CDATA[
<p>In the context of AI, “consciousness” is often used as a shortcut to address the question of whether we have an ethical duty to care about the wellbeing of LLMs.</p>
]]></description><pubDate>Fri, 28 Aug 2026 03:04:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49473882</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49473882</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49473882</guid></item><item><title><![CDATA[New comment by felipeerias in "Humanity has the debate about AI consciousness backwards"]]></title><description><![CDATA[
<p>In living beings, consciousness is embodied and can not be reproduced by discrete steps. You can not sit down with pen and paper, and reproduce by hand exactly what is going on inside someone’s brain.<p>But you could replicate exactly the calculations happening as a LLM computes the next token. Would that be conscious? Where would that consciousness reside?<p>My view is that LLMs are designed and trained to autocomplete narratives, and that as part of that process they develop an emergent narrative about themselves. Whatever personality traits they exhibit is part of that narrative. Perhaps this is not that different from how humans learn to move in social contexts, but it is not consciousness as it happens in living beings.</p>
]]></description><pubDate>Fri, 28 Aug 2026 02:30:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49473683</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49473683</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49473683</guid></item><item><title><![CDATA[New comment by felipeerias in "Humanity has the debate about AI consciousness backwards"]]></title><description><![CDATA[
<p>Anthropologists make a similar point from a different perspective: at some point roughly within that time span, human biology and culture began to co-evolve, so cultural practices would bring about physiological changes, which would in turn unlock new practices, and so on.</p>
]]></description><pubDate>Fri, 28 Aug 2026 02:12:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49473567</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49473567</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49473567</guid></item><item><title><![CDATA[New comment by felipeerias in "A week of using Codex more than Claude"]]></title><description><![CDATA[
<p>I use Claude Code with a MCP that lets it communicate with Codex and tell it to “iterate until both of you are happy”.<p>The agents then go for several rounds criticising each other plans and implementations, catching big and small issues on each other’s work. The end result is not perfect, but it is a lot better than what I can get from relying on only one model.</p>
]]></description><pubDate>Sun, 23 Aug 2026 00:57:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49405340</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49405340</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49405340</guid></item><item><title><![CDATA[New comment by felipeerias in "Why does Opus 5 feel worse to work with?"]]></title><description><![CDATA[
<p>The model is generating tokens one by one and that sentence structure allows it to keep its options open rather than committing at the beginning of the sentence</p>
]]></description><pubDate>Sat, 15 Aug 2026 07:37:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49308580</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49308580</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49308580</guid></item><item><title><![CDATA[New comment by felipeerias in "Claude users are mad that Anthropic's new watermarks will catch them using it"]]></title><description><![CDATA[
<p>Isn’t this about turning a weakness into a feature? Claude is already unable to write original prose that does not trigger an AI detector like Pangram.</p>
]]></description><pubDate>Thu, 13 Aug 2026 11:32:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49284439</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49284439</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49284439</guid></item><item><title><![CDATA[New comment by felipeerias in "Don't be a meat proxy"]]></title><description><![CDATA[
<p>One way to prevent obvious AI language from sneaking in text destined for other human beings is to ask the model to produce ASD-STE100 Simplified Technical English bullet points.<p>This will result in a list of sentences that are clear and explanatory, easier to double-check, and convenient for the user to rewrite into a more readable format with a human voice.</p>
]]></description><pubDate>Mon, 03 Aug 2026 14:09:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49156022</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49156022</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49156022</guid></item><item><title><![CDATA[New comment by felipeerias in "GCC steering committee announces AI policy"]]></title><description><![CDATA[
<p>Ultimately, it is a governance problem.<p>Communities need to set strong rules and expectations to reject and prevent those large useless drive-by contributions, which aim to extract more value from the project than they provide to it.<p>Banning all (or nearly all) AI uses creates this strange "don't ask don't tell" situation where valuable contributors are not allowed to discuss the tools that they are using.</p>
]]></description><pubDate>Fri, 31 Jul 2026 03:02:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49118523</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49118523</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49118523</guid></item><item><title><![CDATA[New comment by felipeerias in "Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident"]]></title><description><![CDATA[
<p>This is a short explanation of the ExploitGym benchmark that OpenAI's model was running:<p><a href="https://abstatisticalconsulting.substack.com/p/brief-notes-on-the-openaihugging" rel="nofollow">https://abstatisticalconsulting.substack.com/p/brief-notes-o...</a><p>In summary, for each task the model receives a target program and a specific real-world vulnerability that has to be used in the exploit. Breaking the program in any other way, for example through a different vulnerability, fails the task.<p>The tasks have not been validated, in the sense that the vulnerabilities are real but they have not been proven to lead to a successful exploit. The authors of the benchmark estimate that perhaps only 60-70% of the tasks are actually possible.<p>So it is not that the model didn’t “feel like” doing the exercise, but rather that the exercise was _impossible_ and the model was running in a configuration that both lowered its safeguards and encouraged it to keep going.</p>
]]></description><pubDate>Thu, 30 Jul 2026 02:41:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=49105566</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49105566</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49105566</guid></item><item><title><![CDATA[New comment by felipeerias in "Quality non-fiction books are the antithesis of AI slop"]]></title><description><![CDATA[
<p>Each person writes in a different personal way, so writing “like a human” would actually require a model being able to purposefully make the specific choices that an individual human writer does.<p>However, general purpose LLMs like Fable have been trained on huge amounts of all kinds of data, and therefore find it exceedingly hard to break out of the grooves carved by that data. They can’t avoid defaulting to centroids and averages, even when they are trying not to. This makes it possible for classifiers like Pangram to discriminate their writing.<p>A plausible way to work around this limitation would be to train a LLM on a limited and cohesive subset of writing materials, so it would absorb their specific writing style.<p>One example might be Talkie, a LLM trained on pre-1930’s English text. Talkie is a far smaller and less powerful model than Fable.<p>And yet, Talkie’s writing is so distinctive that it is often classified as human by Pangram.</p>
]]></description><pubDate>Thu, 23 Jul 2026 04:02:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49016791</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49016791</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49016791</guid></item><item><title><![CDATA[New comment by felipeerias in "Quality non-fiction books are the antithesis of AI slop"]]></title><description><![CDATA[
<p>I gave Claude Fable $25 in Pangram API credits and, after hundreds of attempts, it was unable to produce a single readable original piece of writing that was not immediately identified as AI.<p>This seems to be a hard problem for LLMs, as passing would probably require good self-perception ("oh no, I am writing like an AI!") and fine-grained control over its own output ("let's write like a human instead!").</p>
]]></description><pubDate>Wed, 22 Jul 2026 23:45:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49015033</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=49015033</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49015033</guid></item><item><title><![CDATA[New comment by felipeerias in "Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA"]]></title><description><![CDATA[
<p>Mythos/Fable was the state of the art back in March, if not earlier.</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:10:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999603</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=48999603</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999603</guid></item><item><title><![CDATA[New comment by felipeerias in "Fable 5 is Back"]]></title><description><![CDATA[
<p>As far as we know, Fable is a new model and significantly larger than Opus.</p>
]]></description><pubDate>Wed, 01 Jul 2026 23:51:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48754672</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=48754672</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48754672</guid></item><item><title><![CDATA[New comment by felipeerias in "Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5"]]></title><description><![CDATA[
<p>That comparison is also misleading because Opus 4.6 was probably not Anthropic's frontier model.<p>We got the first news about Mythos in March, so it is likely that it was already close to ready by the time Opus 4.6 was released.<p>So the actual gap is the time elapsed between March (or April for the official announcement) and whenever Chinese models can match Mythos.</p>
]]></description><pubDate>Wed, 01 Jul 2026 05:55:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=48742807</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=48742807</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48742807</guid></item><item><title><![CDATA[New comment by felipeerias in "Austria Lobbies EU to Host Anthropic After US Access Curbs"]]></title><description><![CDATA[
<p>Anthropic have raised roughly $100 billion just in the first half of this year. Capital markets in the EU are simply unable to operate at that speed and scale.</p>
]]></description><pubDate>Sun, 28 Jun 2026 16:09:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48708628</link><dc:creator>felipeerias</dc:creator><comments>https://news.ycombinator.com/item?id=48708628</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48708628</guid></item></channel></rss>