<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: Topfi</title><link>https://news.ycombinator.com/user?id=Topfi</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 19 Aug 2026 01:19:45 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=Topfi" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[YouTube: A view counts from the first frame]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.heise.de/en/news/YouTube-A-view-counts-from-the-first-frame-11416784.html">https://www.heise.de/en/news/YouTube-A-view-counts-from-the-first-frame-11416784.html</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49342707">https://news.ycombinator.com/item?id=49342707</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Tue, 18 Aug 2026 07:49:31 +0000</pubDate><link>https://www.heise.de/en/news/YouTube-A-view-counts-from-the-first-frame-11416784.html</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49342707</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49342707</guid></item><item><title><![CDATA[New comment by Topfi in "Ask HN: Alternatives to GitHub"]]></title><description><![CDATA[
<p>Just signed up, so not much to report yet, but one thing I noticed is that, because the sign in is on tngl.sh rather than a subdomain of tangled.org, it might be confusing for some users, especially as the sign in page has a very different design, theme and layout to the rest of your website and tangled as a whole. Additionally, my password manager didn't show the just created password due to the URL mismatch. Perhaps this could be addressed to improve the experience.<p>Edit: I really like that one-click "watch logs via SSH" copy button. Maybe those could be attached to the top of the top, as it stands the pipeline just pushes them further and further down. Find the UI and UX overall very pleasant.</p>
]]></description><pubDate>Mon, 17 Aug 2026 21:11:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49337701</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49337701</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49337701</guid></item><item><title><![CDATA[GPT-5.6 Sol Pricing Cut by 50% on OpenRouter]]></title><description><![CDATA[
<p>Article URL: <a href="https://openrouter.ai/openai/gpt-5.6-sol">https://openrouter.ai/openai/gpt-5.6-sol</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49337602">https://news.ycombinator.com/item?id=49337602</a></p>
<p>Points: 617</p>
<p># Comments: 442</p>
]]></description><pubDate>Mon, 17 Aug 2026 21:03:18 +0000</pubDate><link>https://openrouter.ai/openai/gpt-5.6-sol</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49337602</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49337602</guid></item><item><title><![CDATA[New comment by Topfi in "Logical – My Codes Are Perfect"]]></title><description><![CDATA[
<p>Given the outage and, having just looked through the Microsoft entry in Firefoxs entitieslist, I felt like sharing a link from a more simple time, before the acquisition of Github.</p>
]]></description><pubDate>Mon, 17 Aug 2026 15:10:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49332329</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49332329</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49332329</guid></item><item><title><![CDATA[Logical – My Codes Are Perfect]]></title><description><![CDATA[
<p>Article URL: <a href="http://logicalawesome.com">http://logicalawesome.com</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49332328">https://news.ycombinator.com/item?id=49332328</a></p>
<p>Points: 1</p>
<p># Comments: 1</p>
]]></description><pubDate>Mon, 17 Aug 2026 15:10:47 +0000</pubDate><link>http://logicalawesome.com</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49332328</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49332328</guid></item><item><title><![CDATA[New comment by Topfi in "Tell HN: Cloudflare silently injects its analytics when you switch nameservers"]]></title><description><![CDATA[
<p>Yeah, agree, purely from a user expectation perspective, strict sounds like it'd suppress everything, not just third-party, especially with lower compatibility turned on.<p>Yeah "browser.events.data.microsoft.com" is in, but mainly because it just ends up under microsoft.com anyways, subdomains are accepted inside the entitieslist.<p>Here the full copy (initially wanted to share via Pastebin but some links triggered a spam filter): <a href="https://cdn.jsdelivr.net/gh/mozilla-services/shavar-prod-lists@e3dc3016f654286992ad51275585580e748bd379/disconnect-entitylist.json" rel="nofollow">https://cdn.jsdelivr.net/gh/mozilla-services/shavar-prod-lis...</a><p>If you or anyone else reading is interested, Mozilla has written some pleasant docs, lots to learn, easy to understand, comprehensive. Could not even consider what I am attempting without the resources they've provided: <a href="https://firefox-source-docs.mozilla.org/toolkit/components/antitracking/anti-tracking/tracking-lists/index.html" rel="nofollow">https://firefox-source-docs.mozilla.org/toolkit/components/a...</a></p>
]]></description><pubDate>Mon, 17 Aug 2026 15:03:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=49332219</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49332219</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49332219</guid></item><item><title><![CDATA[New comment by Topfi in "Tell HN: Cloudflare silently injects its analytics when you switch nameservers"]]></title><description><![CDATA[
<p>Just saw the edit, all clicked now. As described this is expected behavior. The entitieslist (can't link right now because Github is down but got a local copy for some experiments) contains exceptions for owners of tracking URLs, in this case as a resource for Cloudflare.com and others owned by them only. Basically, because they are the same entity, they are considered one. Whether that could be communicated better by upstream, that's worth a discussion. Anyone besides Cloudflare.com has cloudflareinsights blocked.<p>Here the specific entry for context from my local copy of ESR 153:<p>{<p><pre><code>  "entities": {

    "Cloudflare": {

      "properties": [

        "cloudflare-quic.com",

        "cloudflare.com",

        "cloudflare.tv",

        "cloudflarestatus.com",

        "cloudflareworkers.com"

      ],

      "resources": [

        "cloudflare.com",

        "cloudflareinsights.com",

        "cloudflarestream.com"

      ]

    }

  }
</code></pre>
}<p>On a side note, spent a while learning how upstream Firefox works in-depth over the last few months, if ETP didn't block cloudflareinsights on pages outside Cloudflare.com I'd have lost any confidence build up and my project would likely linger even longer. Might just add a setting that totally excludes such exceptions if I can properly test it before release, seems there might be demand. Admittedly more for UX and honest communication clarity reasons (top setting truly prevents everything) then privacy, not the main goal of Hominis as a project.</p>
]]></description><pubDate>Mon, 17 Aug 2026 14:32:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=49331658</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49331658</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49331658</guid></item><item><title><![CDATA[New comment by Topfi in "Incident with Github.com"]]></title><description><![CDATA[
<p>Github being under the CoreAI division probably also doesn't help the engineers prioritize addressing infrastructure issues and makes using LLM load as an excuse feel self-inflicted. Akin to feeling sorry when a pyromaniacs house burns down...</p>
]]></description><pubDate>Mon, 17 Aug 2026 14:08:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49331226</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49331226</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49331226</guid></item><item><title><![CDATA[New comment by Topfi in "Incident with Github.com"]]></title><description><![CDATA[
<p>No one ever has to commit anything time sensitive.</p>
]]></description><pubDate>Mon, 17 Aug 2026 14:03:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49331132</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49331132</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49331132</guid></item><item><title><![CDATA[New comment by Topfi in "Tell HN: Cloudflare silently injects its analytics when you switch nameservers"]]></title><description><![CDATA[
<p>Standard, Strict or Custom? Strict should block that one for sure, will check once at home.</p>
]]></description><pubDate>Mon, 17 Aug 2026 12:51:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49330025</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49330025</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49330025</guid></item><item><title><![CDATA[New comment by Topfi in "Anthropic's 'watermark' text adulteration in Claude is a perversion of writing"]]></title><description><![CDATA[
<p>Fair point, for text it is far harder to prevent signatures being applied to generated text vs images at the moment of capture and most approaches I can come up with to remedy this can either be bypassed (edit histories can be output by models similar to humans) or will be controversial. Taking a page out of the anti-cheat textbook, mainly written for gaming, there are methods which might hold in the medium term. Less a fan of kernel level myself, though it might be worth exploring as there has been massive investment by the games industry into making it somewhat robust, but the approach Valve has taken with VACnet could be an inspiration worth exploring that is less invasive into peoples systems. Keystroke analysis, etc. could be relied upon as a basis for signatures, harder to spoof for current day LLMs over generating edit histories.<p>I will fully admit that at a point in the future, maybe not too soon, models may be trained to bypass that too, at which point we are back where we started. As a skeptic of the extend that capabilities are emergent in LLMs vs specific to training data, I am somewhat hopeful that unless models are specifically trained for evading such human detection solutions, they'd struggle to do so, but it could still end up as a byproduct of improved, lower latency computer use focused training. Not emergent as the term is used in regard to models because that is still output performance improvements clearly traceable to very specific training data, but incidental as the goal of said training data was not to bypass.<p>For what it's worth, I find human authorship being verifiable to simply be the more crucial problem over watermarking model output, so if research is to focus on one, I'd rather it the former. Maybe both signing human authored content and watermarking LLM output are both only possible in the near term, I hope not but fear it that might be the case. If so, we as a society will have some major challenges ahead (beyond all the ones we'd have anyways).<p>Alternatively, we could also just start scanning everyones eyeballs...</p>
]]></description><pubDate>Mon, 17 Aug 2026 12:14:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49329624</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49329624</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49329624</guid></item><item><title><![CDATA[New comment by Topfi in "Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing"]]></title><description><![CDATA[
<p>Why not switch it around? Solves problems of privacy (no data upload), is far more doable and it's far more important to be able to verify content has not been tampered with after originating from a human or capture device (like a camera). We also have decades of cryptographic experience in reliably signing text, images, video, etc. and it doesn't break apart because a text is too short.<p>Having proof that content (especially images and video evidence) is unmodified (whether via Photoshop, Paint or a model) is far more valuable then having evidence that an image was manipulated or generated fully by a model (which still leaves other forms of manipulation), I feel the same goes for human authored vs generated text. Free to admit that using models to generate any kind of media whole-cloth is still unappealing to me and I still pay for commissioned artwork or make it with my limited abilities for what that's worth. Do like to (poorly) write my musings too and see UX as something were thoughtful contributors (like the opinionated, sometimes controversial, but certainly talented GNOME Gitlab contributors) can make a major impact.<p>Code can be beautiful, interesting and serve purpose beyond execution, of course, but for most people, in most cases, it does not in the same way as audiovisual content (not limited to art). Having code just to execute and resolve a problem can have value all in itself, the code being a means to an end whose quality, let us be honest, was barely a concern in most corporations long before LLMs.<p>Also have rarely (honestly never) before LLMs fully owned all parts of any code base, always relied in part on someone's prior effort in (Flutter/Dart mostly) packages, whereas when writing, drawing, etc. I have far more situations where I make something from scratch and everything there is only there because of my conscious decision. Even simple marketing mockups that, quality wise, any modern model would beat feel different when I was fully in control, where to place what, etc. Objectively worse (at my skill level), probably, but still never the same.<p>Knowing something was made from scratch by a human has value to me, beyond misinformation prevention. Knowing for a fact that LLMs were used instead of importing a library, using a template, or something similar that leads to expending similar amounts of effort, I don't see that being nearly as valuable. Heck, with all the importing and my experience back then vs now, I am spending more effort actually fully reading any LLM output in my code then I spent back then auditing Flutter/Dart packages. Then again, LLM output fails far more unpredictable then those messy packages that simply got Gradle to take down my system...<p>Happy to admit, I have been skeptical of watermarking LLM output being feasible for quite some time and having looked into SynthID Text and proposals being researched, I am convinced that it is challenging to impossible beyond the lowest common denominator and less important then proofing human authorship.<p>It will catch people just copying LLM output into their replies without thought, which is not a negative in my book, especially if it is not discernibly affecting output quality in regular use cases. Anyone who wouldn't copy Wikipedia into their dissertation will, in my opinion, be able to bypass text watermarking as proposed however, I feel we need to be honest there.<p>Thing is, if that's the case and text watermarking will only ever catch LLM created slop, is that a bigger problem then the misinformation, harm to creators due to authorship questions and accusations, making it harder to use evidence in proceedings, teachers not trusting students even when they did the work themselves, etc.? Signatures for all such cases will be difficult to implement, yes, but I feel are going to be of greater value in the not to distant future and I equally feel are not impossible, not least because idiots will always want to hide their LLM usage, whereas human authorship is something they take pride in and want to proof.</p>
]]></description><pubDate>Mon, 17 Aug 2026 10:40:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49328873</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49328873</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49328873</guid></item><item><title><![CDATA[New comment by Topfi in "Research papers using "kidney disappointment" instead of "kidney failure""]]></title><description><![CDATA[
<p>I had sent this an hour ago as a quote with a few others (lactose bigotry, of course the kidney disappointment) to someone via iMessage. The Apple Intelligence summary was unfortunately not screenshotted by them, but I am told it was a doozy...</p>
]]></description><pubDate>Sun, 16 Aug 2026 15:55:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49321180</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49321180</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49321180</guid></item><item><title><![CDATA[New comment by Topfi in "What happens when an LLM never sees material beyond fifth grade?"]]></title><description><![CDATA[
<p>> This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.<p>I did say "no reliable internalised way to assess the accuracy of their output", which is something very specific. If that Sonnet output is reasoning traces, there are many issues with using that as a source:<p>For one, back and forth reasoning does often correlate with less, not more accurate overall outputs in evaluations and one example could never seriously be extrapolated to be considered "reliable", i.e. happening consistently and dependably.<p>Secondly, self-correction is also not self-verification in regard to model output, revisions not necessarily mean internal accuracy assessment by themselves (again, over-revisioning has lead LLMs in my and even public evals like the one linked above to step away from accurate information written in their reasoning traces but discounted in the final output (if we must use anecdotal examples like your Sonnet quote)).<p>Then there is the fact that, unless that reasoning trace (if it indeed is one) was copied from a months old chat history, Anthropic has obfuscated their reasoning traces so this output is (if it isn't an ancient history you dug up) from the obfuscation model in between and not reflective of the actual models reasoning. So even for anecdotal evidence, this can likely not be used (unless again, you went for December 2025 history). If this was not reasoning traces (I am fairly confident it is having spend a long time reading Anthropic summarisations vs actual reasoning traces when the switch was on-going and you could get both for a limited time, what you shared reads very much like obfuscated rather than pre-obfuscation reasoning output), that still leaves this as anecdotal, one time, possibly erroneous (did this back and forth even prevent an error in the final output) and there are more issues still with just using that quote as evidence, this is simply unsuitable as a source in any situation.<p>Here are some papers I read lately, all published in 2026 and using the current crop of models which were what led me to make that specific statement. LLMs currently have no reliable internalised way to assess the accuracy of their output, at least as far as the literature is concerned:<p>> Even state-of-the-art models struggle to reliably discriminate between data uncertainty and model uncertainty.<p>Beyond “I Don’t Know”: Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty [0]<p>> LLMs cannot reliably revise their own errors without an external signal.<p>> The same models that confidently catch and repair errors in external content routinely fail to identify identical errors in their own reasoning traces.<p>The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models [1]<p>Simply, as of today, LLMs cannot reliably translate whatever internal signals they have into accurate self-assessment of their outputs. Even in papers that show limited, edge case capability, this often breaks with minimal prompt and/or task changes (happy to link those too, I just need to get to my Mac where I have the PDFs), so it is not reliable.<p>If you got a paper that shows that I am false, that there is reliable, internal self-verification of a models output (even a small research LLM not yet publicly released), I'd be happy to read it and retract my original statement.<p>[0] <a href="https://arxiv.org/html/2604.17293" rel="nofollow">https://arxiv.org/html/2604.17293</a><p>[1] <a href="https://arxiv.org/html/2606.05976" rel="nofollow">https://arxiv.org/html/2606.05976</a></p>
]]></description><pubDate>Sun, 16 Aug 2026 15:20:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49320898</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49320898</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49320898</guid></item><item><title><![CDATA[New comment by Topfi in "What happens when an LLM never sees material beyond fifth grade?"]]></title><description><![CDATA[
<p>Greatly varies by the model, but such a flat denial along with calling those with (verifiably accurate [0]) experiences which happen to be different from yours "the anti-AI crew" and just assuming they can't possible use LLMs and couldn't possibly have formed their opinions on evidence, well, says a lot.<p>LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output and they have, do and will continue to for the foreseeable future, lead people down paths they shouldn't [1], partly by their architecture and the limits of the technology, partly by an intent for maximising retention.<p>For what it's worth on people, I have, both in politics and business, unfortunately made the painful discovery that, whether intentional (because the powerful person in question cannot or doesn't want to deal with different opinions to their own) or unintentionally (because those with sycophantic tendencies simply managed to manipulate themselves into their inner circle over years and became trusted), there are people in (financial, political or other types of) power which are surrounded by few willing to tell them when they are wrong and even fewer that are actually listened to if push comes to shove, which often affects the personality and mental health of said powerful people negatively, to the detriment of society, their family, their employees, etc.<p>Members of the media are actually complicity. I very much disagree that the media presents obscenely powerful people in too negative a light, more the opposite. Am very firm that, to retain access to the rich, powerful and famous, there is far too much sane-washing of utterly ridiculous, unacceptable, harmful and/or dangerous behaviour. Sometimes the person in questions own health and safety are put at risk, because neither the people around them, nor public opinion or reporting treat their behaviour in the way appropriate. Sometimes this again leads to harm for the public, their family, those working under them, etc.<p>What'd get someone with less power or in a lower tax bracket ridiculed or even sectioned is often reported as "eccentricities", just being "passionate about a topic" and trying to do the "marketing rounds".<p>[0] <a href="https://artificialanalysis.ai/?omniscience=omniscience-hallucination-rate#omniscience-tabs" rel="nofollow">https://artificialanalysis.ai/?omniscience=omniscience-hallu...</a><p>[1] <a href="https://arxiv.org/html/2602.19141v1" rel="nofollow">https://arxiv.org/html/2602.19141v1</a></p>
]]></description><pubDate>Sun, 16 Aug 2026 12:56:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49319620</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49319620</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49319620</guid></item><item><title><![CDATA[New comment by Topfi in "Semaglutide linked to lower predicted dementia risk"]]></title><description><![CDATA[
<p>> [...] not make it a bit more explicitly known.<p>Honest question, how could it be made more explicitly known? It's literally under ACKNOWLEDGMENTS, in the most clear language possible.</p>
]]></description><pubDate>Sat, 15 Aug 2026 23:42:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49315367</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49315367</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49315367</guid></item><item><title><![CDATA[New comment by Topfi in "Qwen 3.8 27B"]]></title><description><![CDATA[
<p>Deepseek v4 Pro or GLM 5.3 for software architecture I feel are only deployed for general, rather doubtful those are for narrow stuff, if we are sticking with this threads mentioned models. For small models, sure, narrowly targeted sets which can be self evaluated are amazingly valuable, but I feel beyond 500B we are in a different dimension. Rating any model the size of GLM or V4 Pro in hours I doubt is done beyond pure vibes.</p>
]]></description><pubDate>Sat, 15 Aug 2026 14:09:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49310730</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49310730</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49310730</guid></item><item><title><![CDATA[New comment by Topfi in "Qwen 3.8 27B"]]></title><description><![CDATA[
<p>Anything from ML pipelines for language specific pruning over a Rust/JS/CSS mix codebase to assistance in motorcycle maintenance and different canvas coatings. Most of my evals build on those requirements and especially past failures, whether in pure information, task execution and coding or tool calling beyond the overfitted mainstream. All stuff derived from actual failures encountered, some still only few models come even close to passing. With such a mix, it just takes a while to get any serious opinion on a model. Doubt anyone can do that in such short time, unless their tasks are so simple that most modern models not only succeed but could themselves accurately rate output. If even Fable still confuses PU or wax coated cotton canvas with a nylon shell, or tells me with a straight face to adjust the valves on a bike that has hydraulic lifters that needs experience for human assessment and the time that comes with it. Anyone with less knowledge either wouldn’t see the mistakes staring them in the face and just go by vibes, any model rating these equally can not tell what is accurate and will just go by the output sounding accurate over being. Gives sometimes very interesting results far different to public benchmarks. Inkling, e.g. is more accurate in not telling you to adjust valves that are simply not adjustable then Fable or Sol, which just tell you to adjust every 5000km. Sometimes even when their reasoning and search includes sections about the fact this is not necessary or possible. The beauty of overfitting and unbalanced training data…</p>
]]></description><pubDate>Fri, 14 Aug 2026 19:29:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49303520</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49303520</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49303520</guid></item><item><title><![CDATA[New comment by Topfi in "Qwen 3.8 27B"]]></title><description><![CDATA[
<p>Honest question, how do you assess models this quickly? What metrics are you using? Would love to get my suite from multiple days and hundreds of prompts down to minutes. Got a few first pass tasks I run upon release for an initial experience, but those only work because even Fable and Sol fail despite objectively correct solutions existing, so it works because most models fail, but then, those are consciously not enough for coding, tool use, adherence or task specific inference and assessment…</p>
]]></description><pubDate>Fri, 14 Aug 2026 16:41:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49301193</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49301193</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49301193</guid></item><item><title><![CDATA[New comment by Topfi in "Choosing an AI model: one prompt, 11 models, different results"]]></title><description><![CDATA[
<p>Sure, but you then only are on Google Maps, Search and other services by them. Have your own website and be indexed by everyone from Kagi, over Bing to DDG.</p>
]]></description><pubDate>Fri, 14 Aug 2026 08:41:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49296114</link><dc:creator>Topfi</dc:creator><comments>https://news.ycombinator.com/item?id=49296114</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49296114</guid></item></channel></rss>