<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: frumiousirc</title><link>https://news.ycombinator.com/user?id=frumiousirc</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 18 Aug 2026 15:04:10 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=frumiousirc" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by frumiousirc in "The Benchmarkpocalypse"]]></title><description><![CDATA[
<p>That is not what I read from danluu's words.  He merely stated in the prompt that there is a holdout set and did not iterate to minimize error against the holdout set.  In a prior attempt he prompted with only "don't overfit" to ill effect on the holdout eval.  Did I misread?</p>
]]></description><pubDate>Tue, 18 Aug 2026 10:36:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49343707</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49343707</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49343707</guid></item><item><title><![CDATA[New comment by frumiousirc in "Qwen 3.8 27B"]]></title><description><![CDATA[
<p>> Bicycle is the right shape. Pelican beak is excellent. Nice background.<p>Relative to other results I agree. But on an absolute measure, there is not a single element in the current bicycle that is real-world accurate and many elements are omitted or non-physical (eg, the transparent seat tube top, entire lack of a head tube).<p>Consider a series of followup benchmarks.<p>With a fresh context of the LLM under test, ask it to generate a list of findings for how the pelican-on-a-bicycle SVG that was produced is inaccurate compared what the real world scene might appear, accepting for the limitations of SVG as a medium.  Then, feed back the list of findings to the original context for a second try. The benchmark can stop here by humans looking at the result and forming their own conclusion.<p>Next phase is to repeat the analysis phase using the 2nd context to determine what findings were satisfied and what new inaccuracies are found.  These two differences can form a second benchmark.<p>Last phase is to iterate with the goal to drive the number of findings to zero.</p>
]]></description><pubDate>Sat, 15 Aug 2026 11:15:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49309637</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49309637</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49309637</guid></item><item><title><![CDATA[New comment by frumiousirc in "Auto mode is now the default in Claude Code"]]></title><description><![CDATA[
<p>I use bubblewrap, which I believe claude code also has internally but not for its `Bash()` tool.<p>I wrap bubblewrap in a script that supports config files to allow different "profiles" of use (analogous to eg firefox profiles).  The bwrap starts with the whole filesystem mounted read-only, then mounts the current directory read-write and then applies further bind mounts for devices, special case other read-write (eg, ~/.cache/) and to mount empties to cover sensitive directories (eg, ~/.ssh/).  The profile also specifies the default command to run and for claude, it gets yolo mode.</p>
]]></description><pubDate>Mon, 10 Aug 2026 11:05:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49242120</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49242120</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49242120</guid></item><item><title><![CDATA[New comment by frumiousirc in "U.S. Department of Energy Launches the Genesis Open Models Initiative"]]></title><description><![CDATA[
<p>Your question piqued my curiosity.<p>I thought maybe NERSC doesn't accept jobs from private corporations but I checked and that's not true, as long as results are not held proprietary.<p>Perhaps ANL was used as they have a lot of compute and they lead and host the Genesis Open Models Initiative?<p>From your link, 2048 B300 GPUs were used for 6 months.  If google search is right, NERSC has 7168 A100.  B300's are way more capable than A100.  To do this training in 6 months, "3 NERSCs" would be needed.<p>Between the political angle and the technical, I'd guess these two make up a big chunk of the answer.</p>
]]></description><pubDate>Sun, 09 Aug 2026 10:44:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49230214</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49230214</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49230214</guid></item><item><title><![CDATA[New comment by frumiousirc in "A Tome of Forbidden Technologies"]]></title><description><![CDATA[
<p>Well, vinyl records are an analog medium so their sampling rate is infinite.  More important would be physical and electrical band limits of the original recording, cutting and pressing devices and of course the playback devices.</p>
]]></description><pubDate>Sat, 08 Aug 2026 12:38:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49221276</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49221276</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49221276</guid></item><item><title><![CDATA[New comment by frumiousirc in "U.S. Department of Energy Launches the Genesis Open Models Initiative"]]></title><description><![CDATA[
<p>There's no mention of "LLM" nor "language".  It does mention "foundation model" which includes LLMs but that also includes non-LLM architectures and non-text data.  Many of the Genesis Initiative proposals answer "foundation model" call with non-LLM systems.  All the FM's I know about currently in this sphere are non-LLMs.  The "about gs1" page also does not mention "LLM" but does talk more about agentic harness and workflows. That description certainly sounds LLM'ish but describes a more rich system.  I don't mean to suggest that LLMs will not be part of these "genesis open models" but as described, this will not result in a replacement for the "claude" or "codex" commands.</p>
]]></description><pubDate>Sat, 08 Aug 2026 11:14:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49220718</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49220718</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49220718</guid></item><item><title><![CDATA[New comment by frumiousirc in "U.S. Department of Energy Launches the Genesis Open Models Initiative"]]></title><description><![CDATA[
<p>> and there are career civil servants (all the government scientists are under this category)<p>US national lab scientists are not even civil servants.  The labs themselves are run by a corporation under contract to the DOE and the scientists work for that corp.  The managing corporation changes from time to time and the scientists transparently start working for whatever assumes the replacement.  The land, the hardware, the buildings and any physical products are owned by the US gov't.  To a very large extent, the intellectual output is set free to the world in the form of papers, presentations and to some small extent (eg compared to CERN) in the form of software.</p>
]]></description><pubDate>Sat, 08 Aug 2026 11:00:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49220630</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49220630</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49220630</guid></item><item><title><![CDATA[New comment by frumiousirc in "Herdr is joining Y Combinator. The runtime stays open"]]></title><description><![CDATA[
<p>> I find myself to be a huge bottleneck, having to do technical design and design reviews, to make sure they're actually working on the right things.<p>I find this depends on the LLMs being used and the person using them and the problems being solved.<p>Like, Fable can be sent off to do some big thing for a long while and I find myself getting bored and move over to push on some other thing.  Or, while reading a paper I have a string of follow up questions and ideas and launch them via web or CLI agent.  Or I bounce between multiple chores in different packages, each of which is fast for an LLM and a low cognitive burden for me.  Then other times, something needs my full, ongoing and serial attention where the LLM turn is only a minor element. I'll use the brief LLM interludes to get up, walk around a bit, stretch and think.<p>As for Herdr, it seems nice when I tried it a few weeks ago.  But, I'm surprised it is getting so much attention (kudos).  Like many people, I've developed more than one work-alike before Herdr hit the scene.  They were based on tmux which I have concluded I simply detest.  In the end, I made a little agent hook script that speaks to Kitty terminal and that plus Kitty and a stable SSH persistent connection gives 80% of what I was looking for.  I do still like the "control panel" aspect of my past and herdr's tries.  Adding that would get me another 10%.</p>
]]></description><pubDate>Fri, 07 Aug 2026 11:46:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49208964</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49208964</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49208964</guid></item><item><title><![CDATA[New comment by frumiousirc in "Prevent cognitive debt by manually retyping LLM-generated code"]]></title><description><![CDATA[
<p>Of course I push my car to the store!  It has the GPS and I've lost the ability to self-navigate.  Plus, how else could I carry my groceries.<p>:)</p>
]]></description><pubDate>Tue, 04 Aug 2026 09:52:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49166333</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49166333</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49166333</guid></item><item><title><![CDATA[New comment by frumiousirc in "Prevent cognitive debt by manually retyping LLM-generated code"]]></title><description><![CDATA[
<p>I think this characterization misses the feedback loop that exists.<p>The OP is not sitting in a one-way flow of info:<p><pre><code>    LLM->human->code 
</code></pre>
but rather a the center of a feedback loop:<p><pre><code>    LLM<-->human<-->code.  
</code></pre>
Human-is-the-loop, not human-in-the-loop.  Each iteration of that loop is fully driven by the human.  Human creativity is involved in both directions.<p>I'd argue that anytime we drive an LLM through more than one turn (and/or more than one session) we are really doing a human-is-the-loop thing.  The OP's extreme case of begin the only thing editing code is on a spectrum with the other extreme being vibe coding (never looking at output code, but still interacting with that output in some way).<p>Off the spectrum is what I call LLM-vomit.  A human one-shots something and puts it out for others to see, suffer and clean up or ignore.  This is code that is encountered literally out of context and can only be further improved (if that is even attempted) by approaching it from first principles.  Such code is akin to people using LLM to generate an answer delivered to another human.  Both flavors (sorry) of LLM-vomit are bad.  I can prompt my own LLM to do that, don't do it for me.</p>
]]></description><pubDate>Tue, 04 Aug 2026 09:50:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=49166317</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49166317</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49166317</guid></item><item><title><![CDATA[New comment by frumiousirc in "PISIGuard: Protect your personal and sensitive info when you chat with AI"]]></title><description><![CDATA[
<p>It could be useful more as a labor saving device to enable censoring while allowing the human to engage in easy copy-paste.  Examples I think of are inclusion of internal DNS names, usernames and human names and email addresses that may be intermixed with log or command output needed to debug some issue.<p>That said, if these strings are really sensitive and I'm too lazy (or not trusting the human user) to self-censor, I'd not rely on this particular kind of tool for censoring.  I'd use local LLMs or cloud LLMs where privacy is part of the contract.</p>
]]></description><pubDate>Mon, 03 Aug 2026 10:36:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49153885</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49153885</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49153885</guid></item><item><title><![CDATA[New comment by frumiousirc in "Gpiozero Flow"]]></title><description><![CDATA[
<p>A key feature about data-flow programming that seems too often missed is that it is (or can be) hierarchical.<p>Define a subgraph of atomic nodes as itself a node with its ports formed from as-yet unconnected ports of its atomic constituents.  Compose yet higher subgraphs of subgraphs and atomic nodes.  Package all this in some way.<p>This is directly analogous to syntactic programming where functions aggregate other function calls and all that packaged into a library with an API.</p>
]]></description><pubDate>Thu, 30 Jul 2026 10:59:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49108264</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49108264</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49108264</guid></item><item><title><![CDATA[New comment by frumiousirc in "GrapheneOS protections against data extraction from locked devices"]]></title><description><![CDATA[
<p>Perhaps bad/naive ideas, but consider these options for unlocking where the requirements changes after one incorrect entry.<p>- After an incorrect key, require two (or more) consecutive valid key entries.<p>- After an incorrect key, no longer accept that nominal key until a secondary key is supplied.</p>
]]></description><pubDate>Mon, 27 Jul 2026 11:54:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=49068268</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49068268</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49068268</guid></item><item><title><![CDATA[New comment by frumiousirc in "You can view a lot of shared conversations via Google"]]></title><description><![CDATA[
<p>It seems Google already /dev/null's this search term.<p><pre><code>    Your search - site:claude.ai/share - did not match any documents.</code></pre></p>
]]></description><pubDate>Mon, 27 Jul 2026 11:37:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49068086</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49068086</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49068086</guid></item><item><title><![CDATA[New comment by frumiousirc in "Ruff v0.16.0 – Significant new updates – 413 default rules up from 59"]]></title><description><![CDATA[
<p>> I really wish Ruff would introduce something similar to Nix’s stateVersion<p>Two approaches:<p>You can pin the version of `ruff` in `pyproject.toml`.<p>You, to let ruff command version advance but pin the settings to a prior version:<p><pre><code>    uvx ruff@0.15.22 check --isolated --show-settings | flat2toml > pinned.toml
    </code></pre>
The `flat2toml` script is left as an exercise.  Then in `ruff.toml`:<p><pre><code>    extend = "pinned.toml"</code></pre></p>
]]></description><pubDate>Mon, 27 Jul 2026 11:04:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49067791</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49067791</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49067791</guid></item><item><title><![CDATA[New comment by frumiousirc in "Ruff v0.16.0 – Significant new updates – 413 default rules up from 59"]]></title><description><![CDATA[
<p>ruff is not new.  It's also truly really fast.  On an as-yet un-ruff'ed 32k SLOC python package it returns in 158ms finding 2k errors.  Having it fix errors is also very fast.<p>What's really tiring is reading such facile "critiques", especially when they are not even applicable.</p>
]]></description><pubDate>Sun, 26 Jul 2026 14:13:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49058443</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49058443</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49058443</guid></item><item><title><![CDATA[New comment by frumiousirc in "Ghost Cut – Or why Cut and Paste is broken everywhere"]]></title><description><![CDATA[
<p>I'd be happy if things stopped interfering with standard X11 select/paste.  It's now a random crap shoot if I have to hold Shift or not to regain that elegance.<p>And, the whole idea of copy/paste having to use both mouse AND keyboard (select Ctrl-c click Ctrl-v) is barbaric.  X11 got it right and it seems all "modern" apps are trying really hard to introduce barbarisms.</p>
]]></description><pubDate>Thu, 23 Jul 2026 11:13:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49019728</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49019728</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49019728</guid></item><item><title><![CDATA[New comment by frumiousirc in "Malleable computing, Emacs, and you"]]></title><description><![CDATA[
<p>> For Markdown to Org translation, Pandoc will be used.<p>Hmm, pandoc?  Surely....<p><pre><code>    M-x org-import-<TAB>
</code></pre>
Damn, why does this family of functions not yet exist?</p>
]]></description><pubDate>Thu, 23 Jul 2026 11:03:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49019625</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49019625</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49019625</guid></item><item><title><![CDATA[New comment by frumiousirc in "Trump administration Announces $5B for the Genesis Mission (AI for Science)"]]></title><description><![CDATA[
<p>I don't know the full break down but some is going to university and DOE lab groups and some of those are even exploring interesting and useful ideas.</p>
]]></description><pubDate>Wed, 22 Jul 2026 19:03:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49011809</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=49011809</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49011809</guid></item><item><title><![CDATA[New comment by frumiousirc in "Mirror your GitHub repos to tangled.org automatically"]]></title><description><![CDATA[
<p>It's fitting a theme of self-deprecation if we consider the origin of the name "git", no?</p>
]]></description><pubDate>Sun, 19 Jul 2026 14:46:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=48968655</link><dc:creator>frumiousirc</dc:creator><comments>https://news.ycombinator.com/item?id=48968655</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48968655</guid></item></channel></rss>