<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: dmrivers</title><link>https://news.ycombinator.com/user?id=dmrivers</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 24 Jul 2026 06:11:02 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=dmrivers" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by dmrivers in "Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong"]]></title><description><![CDATA[
<p>Cool idea, some feedback:<p>1. Consider using conformal prediction to calibrate the cutoff. Conformal prediction provides a distribution-free guarantee under exchangeability. This would let you turn your raw probe score into a threshold with a guaranteed bound on the rate of wrongly-kept on-device answers. Source: <a href="https://en.wikipedia.org/wiki/Conformal_prediction" rel="nofollow">https://en.wikipedia.org/wiki/Conformal_prediction</a><p>2. The best indicators of confidence in ML come from multiple independent methods. What was the result if you combine the token entropy and verbal confidence reporting methods? Does this improve the result?<p>3. I noticed you didn't mention the assessment method of rerunning the model and judging whether outputs are consistent. How does that method compare in terms of AUROC?</p>
]]></description><pubDate>Thu, 23 Jul 2026 10:16:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49019286</link><dc:creator>dmrivers</dc:creator><comments>https://news.ycombinator.com/item?id=49019286</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49019286</guid></item><item><title><![CDATA[New comment by dmrivers in "Making"]]></title><description><![CDATA[
<p>But does this mean that a 99.99% reliable LLM would turn us back into the mode of us building it again? I would not say so.<p>I think for tasks that are about decisions, having the LLM make decisions is what makes it feel like the LLM did something for me.<p>Consider mowing the lawn. Imagine I had a lawnmower robot that does the mowing all on its own. Despite perfect accuracy, I didn't mow the lawn; it did. If I sit on that lawnmower the whole time and start driving it instead of letting it go on its own, then I mowed the lawn. Even if I stand there and control it with a joystick, I still mowed the lawn. Ownership comes from the decisions about where to mow.</p>
]]></description><pubDate>Wed, 22 Jul 2026 17:54:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49010743</link><dc:creator>dmrivers</dc:creator><comments>https://news.ycombinator.com/item?id=49010743</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49010743</guid></item><item><title><![CDATA[New comment by dmrivers in "Back to Kagi"]]></title><description><![CDATA[
<p>I would be interested in seeing that block list if you have it somewhere</p>
]]></description><pubDate>Wed, 22 Jul 2026 13:57:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=49006928</link><dc:creator>dmrivers</dc:creator><comments>https://news.ycombinator.com/item?id=49006928</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49006928</guid></item><item><title><![CDATA[New comment by dmrivers in "Show HN: A comprehensive, filterable list of AI agent jails"]]></title><description><![CDATA[
<p>Interesting. I suggested FreeBSD jails as a PR. Free BSD Jails are great! I use them on <a href="https://www.nearlyfreespeech.net/" rel="nofollow">https://www.nearlyfreespeech.net/</a> and they work well. Classic, long-running jail.</p>
]]></description><pubDate>Mon, 20 Jul 2026 14:09:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48979105</link><dc:creator>dmrivers</dc:creator><comments>https://news.ycombinator.com/item?id=48979105</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48979105</guid></item><item><title><![CDATA[New comment by dmrivers in "Show HN: Leaves – A text-UI disk usage treemap visualizer"]]></title><description><![CDATA[
<p>I found the bottom-right corner of things in my home dir I had no idea about was the lowest hanging fruit. I think this is because my typical cleaning routine was either A. to have claude code find big files I could delete or B. du -sh -- *, which is too slow on ~. Thanks for the decluttering!</p>
]]></description><pubDate>Mon, 20 Jul 2026 13:37:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48978724</link><dc:creator>dmrivers</dc:creator><comments>https://news.ycombinator.com/item?id=48978724</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48978724</guid></item></channel></rss>