<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: 5555watch</title><link>https://news.ycombinator.com/user?id=5555watch</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 03 Sep 2026 08:52:22 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=5555watch" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by 5555watch in "I trained a small transformer in 1.5hrs and it beats many LLMs"]]></title><description><![CDATA[
<p>It means the questions including their answers are dependent. Ie, theres a data generating process for them, that the model uncovers. Like a KNN is known to have near Bayes accuracy as k/n to 0, n to infty, k to infty. The data reveals the dgp.<p>I suspect if you feed unrelated or even garbage questions into the eval set, it would reduce the performance.</p>
]]></description><pubDate>Tue, 01 Sep 2026 22:19:15 +0000</pubDate><link>https://news.ycombinator.com/item?id=49529009</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49529009</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49529009</guid></item><item><title><![CDATA[New comment by 5555watch in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>It makes sense. Even if it finds some exploit on your own code, who's to say you can't reuse the same exploit on some other system?</p>
]]></description><pubDate>Tue, 01 Sep 2026 21:25:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49528457</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49528457</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49528457</guid></item><item><title><![CDATA[New comment by 5555watch in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>While I can't speak for everyone in academia, I personally don't feel comfortable in putting my research questions and outputs to a private website, before the idea is at least arxived. Especially as all the Fable/Mythos prompts are said to be human reviewed.<p>So I believe that, at least in the short run, we might be seeing breakthroughs in hard open problems or in low hanging problems which are not that interesting to spend time on.<p>I may be wrong, if some research labs have private contracted access to the models</p>
]]></description><pubDate>Tue, 01 Sep 2026 20:06:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49527379</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49527379</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49527379</guid></item><item><title><![CDATA[New comment by 5555watch in "Claude Fable 5.1 and Claude Mythos 5.1"]]></title><description><![CDATA[
<p>Is Fable 5.1 still actively downthrottling the reasoning when questions relate to frontier ML questions, like it did with 5.0?</p>
]]></description><pubDate>Tue, 01 Sep 2026 20:00:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49527301</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49527301</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49527301</guid></item><item><title><![CDATA[New comment by 5555watch in "Understanding ChatGPT Work"]]></title><description><![CDATA[
<p>It's not clear to me why isn't the 5.6 Sol Pro available on Codex. Maybe due to the high computational demands and slow speed. But as an anecdote, once when the Sol Ultra models were stumped on one problem, I gave the full extended documentation and all the details to the Sol Pro model to review and provide a suggestion. It eventually did, but the Ultra pushed back against the proposals, and they also seemed sketchy to me as well. We eventually tried them and they didn't work.<p>Maybe the Sol Pro is not that great anymore, considering the token vs output balance.</p>
]]></description><pubDate>Mon, 31 Aug 2026 17:45:52 +0000</pubDate><link>https://news.ycombinator.com/item?id=49512571</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49512571</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49512571</guid></item><item><title><![CDATA[New comment by 5555watch in "What we lost when search stopped making us think"]]></title><description><![CDATA[
<p>I can't be the only one that scrolls past quickly the whatever Google-AI spurts out. If I wanted to ask an AI - I would. But I've learned that they hallucinate stuff with numbers or anything that's time sensitive/recent (even if that might not be the truth now). So I won't spend my time on verifying their output. I'm after a quick search, not a likely guess.<p>What's worse, when reading the Google outputs, I have learned to ignore almost any website besides a select few, one of which is Reddit. Which is very weird thinking about it, but it worked in the past.</p>
]]></description><pubDate>Fri, 21 Aug 2026 15:53:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49389977</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49389977</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49389977</guid></item><item><title><![CDATA[New comment by 5555watch in "Cerebras CS-4"]]></title><description><![CDATA[
<p>>Everyone should occasionally go back to the old models to see how much worse they were, like even a year ago you could generate results but they were typically full of bugs<p>Oh yeah, I'm still amazed how good the current iteration of models are for coding (I have a fear it's too good to be true - so will get taken away..). Exactly a year ago I switched from GPT 5 to Gemini just because the coding with R language was terrible; and even with Python it kept forgetting and mixing basic stuff. Gemini at the time had much longer context window and was miles ahead on R syntax.<p>Current experience of just leaving a Codex Agent chug until a stable solution is completed is still mind blowing to me.</p>
]]></description><pubDate>Wed, 19 Aug 2026 16:55:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=49364076</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49364076</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49364076</guid></item><item><title><![CDATA[New comment by 5555watch in "GPT 5.6 Sol is the best "vision" model OpenAI ever released"]]></title><description><![CDATA[
<p>All of your use cases are very advanced.<p>I recently used it at grocery stores in a foreign country. Photographed the whole aisle and told it to find Y (detergent, softener, glue, sour cream, whatever), at the same time recommend the best Y for whatever reason. Worked marvelously, including the cases where the object wasn't present and it told me there was nothing useful.<p>I asked then, can you crop the exact image of how does the item look like and where is it in the aisle - did that perfectly as well.<p>I will add that all frontier models were fine with such tasks from the early 2024's.</p>
]]></description><pubDate>Mon, 17 Aug 2026 13:22:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49330417</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49330417</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49330417</guid></item><item><title><![CDATA[New comment by 5555watch in "Learning more about Claude's mathematical capabilities"]]></title><description><![CDATA[
<p>But the words don't really matter, do they? The model thought "user said believe in yourself, it means they want me to continue"..<p>AI did a fixed amount of guesses, didn't yield anything. It probably documented the tries, outcome, and some numbers hinting at why they failed. So the user could have prompted "continue", or "try again with previous outcome in mind, generate new ideas and test them" and it would probably yield the same result.</p>
]]></description><pubDate>Tue, 11 Aug 2026 10:38:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49256041</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49256041</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49256041</guid></item><item><title><![CDATA[New comment by 5555watch in "Honey, I shrunk the embeddings: Matryoshka vs. PCA"]]></title><description><![CDATA[
<p>PCA is good, but you could also try playing around with Sparse (robust) PCA. The sparsification loses orthogonality, but does not necessarily lose information, it can yield a different rotation and cleaner vectors. Now whether that matters in the context of LLMs/Embeddings - I cannot tell.</p>
]]></description><pubDate>Sun, 09 Aug 2026 16:57:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49233169</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49233169</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49233169</guid></item><item><title><![CDATA[New comment by 5555watch in "AMD acquires Taalas to boost inference performance by etching models in silicon"]]></title><description><![CDATA[
<p>Even for consumers.. Every enthusiast now wants those specced out rigs to play with LLMs. Make a nice chip for that, and it will reduce some pressure on consumer RAM demand</p>
]]></description><pubDate>Fri, 07 Aug 2026 22:25:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49216961</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49216961</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49216961</guid></item><item><title><![CDATA[New comment by 5555watch in "AMD acquires Taalas to boost inference performance by etching models in silicon"]]></title><description><![CDATA[
<p>Is it still true? I'd assume you should be able to freeze the matrix and unfreeze an expert block, before feeding particularly chosen training data. Or that doesn't work?</p>
]]></description><pubDate>Fri, 07 Aug 2026 21:38:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=49216535</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49216535</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49216535</guid></item><item><title><![CDATA[New comment by 5555watch in "AMD acquires Taalas to boost inference performance by etching models in silicon"]]></title><description><![CDATA[
<p>In my understanding the first Deep Think / Pro models were already very good as they were doing some kind of parallel repeated reasoning, thus were slow and expensive. So if chatjimmy speeds enables a fast deep think level performance, I think that would be great.</p>
]]></description><pubDate>Thu, 06 Aug 2026 23:23:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=49203990</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49203990</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49203990</guid></item><item><title><![CDATA[New comment by 5555watch in "The Dunning-Kruger effect may just be a data artefact (2020)"]]></title><description><![CDATA[
<p>I mean in their code the simulation is incorrect. Someone found the source: <a href="https://github.com/pem725/Dunning-Kruger" rel="nofollow">https://github.com/pem725/Dunning-Kruger</a><p>They're sampling slope and bias for the line from U[0,1] and U[0,100], the expected values of which will be 0.5 and 50. That's why they get what they get.<p>They say that randomness is any positively sloped line with any positive bias. So they assume dependence, just a very high variance of it. It's incorrect as they miss the negative half of the parameters - it would then correctly yield zeroes, for the presumed independence.</p>
]]></description><pubDate>Tue, 04 Aug 2026 10:37:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49166681</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49166681</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49166681</guid></item><item><title><![CDATA[New comment by 5555watch in "The Dunning-Kruger effect may just be a data artefact (2020)"]]></title><description><![CDATA[
<p>It's an interesting read. Curiously, it doesn't really debunk anything.<p>The fact that given X and Y random and independent, that Y-X is correlated with X doesn't disprove the Dunning Kruger. It in fact proves that Y = 1  X is a poor predictor, and the true model is Y = 0  X. In other words, perceived ability (of the human) cannot predict the actual test scores. Which is exactly what DK claims, but to a very extreme effect.<p>Note that, if there is actual signal (plus noise), e.g., if Y = X + eps; so the actual score is exactly the perceived score plus some added variation, the (Y-X)~X will be uncorrelated. In such case, there will be no DK effect, because the users are good at predicting their actual test scores, plus some constant variation.</p>
]]></description><pubDate>Mon, 03 Aug 2026 23:12:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49162583</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49162583</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49162583</guid></item><item><title><![CDATA[New comment by 5555watch in "Ten advances in mathematics and theoretical computer science"]]></title><description><![CDATA[
<p>The extrapolation can also be a learned skill, especially in math. How many papers took result X, extended it to Y using known building blocks, and applied to Z.<p>By the way, convex hull permits extrapolating past the training data. LLM won't invent a new word that could not be defined by a sequence of known words. Just if it's meaningless and fully random/hallucinated, the new knowledge won't work with other known information blocks (breaks convexity).</p>
]]></description><pubDate>Mon, 03 Aug 2026 21:48:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49161916</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49161916</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49161916</guid></item><item><title><![CDATA[New comment by 5555watch in "The Dunning-Kruger effect may just be a data artefact (2020)"]]></title><description><![CDATA[
<p>They're simulating randomness incorrectly: relationship between true and perceived will average 0.5, not 0; and bias will average 50%, not 0%. That's why their "random data" is sloped.<p>Add negative relationship and negative bias, and the random data will act as intended - hovering randomly around 50%.</p>
]]></description><pubDate>Mon, 03 Aug 2026 21:25:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49161641</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49161641</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49161641</guid></item><item><title><![CDATA[New comment by 5555watch in "The Dunning-Kruger effect may just be a data artefact (2020)"]]></title><description><![CDATA[
<p>Thanks for finding the code!<p>Now it's much more clear. The simulated data tries generating the true relationship between actual and perceived scores from 0.0 to 1.0, and bias in self-reporting from 0% to 100%.<p>So the output graph should be the average of all these data generating processes, yielding perceived relationship around 0.5 and bias around 50%, with some high variation.<p>If you have the access, run their Shiny code with these values, and you will see the published plot.<p>I'd argue that this demonstration is much weaker than "making original no more meaningful than random". It's more that the "simulated 50% bias and 0.5 true correlation looks similar to what DK published", which is also far fetched given the data generation they did.<p>Note: true random (what they were going for) would cover negative relationships, yielding the random true relationship around 0; and if they wouldn't correct the sign of Bias, it would also average at around 0; yielding a realistic "random" with the slope hovering about 50% for any percentile.</p>
]]></description><pubDate>Mon, 03 Aug 2026 21:13:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49161519</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49161519</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49161519</guid></item><item><title><![CDATA[New comment by 5555watch in "The Dunning-Kruger effect may just be a data artefact (2020)"]]></title><description><![CDATA[
<p>So what does that change? If there's no error bars on the graphs you can't discuss significant differences easily. And if they do differ, plotting differences will yield the U shape curve.</p>
]]></description><pubDate>Mon, 03 Aug 2026 20:57:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49161344</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49161344</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49161344</guid></item><item><title><![CDATA[New comment by 5555watch in "The Dunning-Kruger effect may just be a data artefact (2020)"]]></title><description><![CDATA[
<p>Very hard to understand the meat behind all the fluff of the article, especially as the simulation code is not available, and as the presented simulated and original graphs are effectively the same (I don't see a disagreement).<p>It's clear that the perceived curve will be differently sloped, as no one will evaluate themselves as the topmost or the bottommost percentiles, so the edges will be biased.<p>And if in both cases we draw differences between perceived and actual, we will get the same curve that everyone knows, biased or not.</p>
]]></description><pubDate>Mon, 03 Aug 2026 20:32:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49161018</link><dc:creator>5555watch</dc:creator><comments>https://news.ycombinator.com/item?id=49161018</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49161018</guid></item></channel></rss>