<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: rafiss</title><link>https://news.ycombinator.com/user?id=rafiss</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sun, 11 Oct 2026 23:40:30 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=rafiss" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>The closest thing we have is if a merged change has to be reverted, the Safety Department holds a Morbidity & Mortality conference and writes up what went wrong. A patient that needs three or more rounds of rework gets one too. There have been seven reverts so far, so not many funerals.</p>
]]></description><pubDate>Sun, 11 Oct 2026 18:14:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=50046079</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50046079</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50046079</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>Sorry about that. The repo is internal to Cockroach Labs for now. It was an issue about adding a progress rollup comment to decomposed parent issues. We’ll get the post updated.</p>
]]></description><pubDate>Sun, 11 Oct 2026 18:12:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=50046058</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50046058</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50046058</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>I like this. Today the parent issue closes automatically once every child is checked off, and nothing re-reads the parent’s original ask against what actually shipped. There is a deferred-scope auditor that catches the deferrals an agent admits to. Your "verify node" would catch the ones it doesn’t, which is the more dangerous kind. I think it would fit naturally as one final child per decomposed parent.</p>
]]></description><pubDate>Sun, 11 Oct 2026 18:03:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=50045974</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50045974</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50045974</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>We don’t run the most expensive model everywhere anymore. We started on Opus for every stage, then reshuffled several times. We landed on Fable doing planning and code review, because code review is the last gate and its bug findings were worth the most. Opus does coding and plan review, and Sonnet runs the nurse roles. The planner can also route mechanical changes to Sonnet. The plan reviewer re-checks that choice, and any rework goes back to Opus.<p>That’s tuning from the cost data, so it's not just totally based on feeling, but it's not a controlled A/B test. I agree that would be worth doing, and I think our experiment here is helping us make the case that it's worth spending more time and effort on running more controlled evaluations.</p>
]]></description><pubDate>Sun, 11 Oct 2026 18:00:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=50045936</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50045936</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50045936</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>I think that’s the real tradeoff. The metaphor brings in behavior we never wrote down, which is both a part of why it works as well as a risk. One reviewer agent declined to send a PR back over a docstring nit “while the patient is exhausted.” It was a fair call, but there's no explicit instructions in any of the prompts/skills that said to do that.<p>We contain it by keeping the persona to a few sentences and writing the rest as explicit scripts or workflows. The rules that really have to hold aren’t left to the role. Human merge approval is a script checking for a real GitHub approval on the PR. A CI failure or a merge conflict bounces the PR before a reviewer agent is even launched. So the metaphor shapes judgment calls, and the hard rules about what should or shouldn't happen are in code.</p>
]]></description><pubDate>Sun, 11 Oct 2026 17:57:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=50045909</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50045909</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50045909</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>I think this is definitely a legit failure mode with this model. We’ve hit some parts of it. For example, a plan reviewer kept pushing a technically valid but negligible security finding through five review rounds and about $70 before a human closed it as won’t-fix. The "Research Department" ran every three hours and was filing around 8 new issues a day, more than we could review. Review nits used to spawn new issues of their own.<p>What helped us maintain a degree of sanity is that every stage records tokens, time, and cost in a ledger, so these show up as numbers rather than vibes. The fixes were mostly about giving a stage permission to stop. Plan review can now propose won’t-fix when the exposure is negligible. Research runs daily, with a cap on unreviewed proposals. Small fixes now ride along in the PR instead of becoming new issues.<p>The other thing is that the quality bar should match the domain. Our migration tools move customer data into a database, so our question was “would we be embarrassed to have merged this?” For small apps I’d agree most of this is overkill.</p>
]]></description><pubDate>Sun, 11 Oct 2026 17:52:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=50045862</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50045862</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50045862</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>To add a bit more: it was motivated by a general desire to not let costs spiral out of control, and when I read the Anthropic context-engineering post I wanted to apply some of the lessons.<p>The audit was agent-driven, but it worked from a rubric I came up with. For example:
- the same rule restated in several places
- narration about what other stages do later
- history and rationale that don’t tell the agent what to do<p>One analyzer agent per skill file flagged passages in each category with a word estimate. Each also produced a “keep” list of things that must not change: commands, gates, templates, sentinels. Agents also did the trimming with some mechanical rules: no command, bash block, label, template, or numeric threshold could change, and we checked the diffs for that. The first pass cut about 18%, because it kept anything it wasn’t sure about. A second pass used an auditor plus an adversarial verifier for each file, and took another ~600 lines out of the six worst files.<p>The more durable result from all of this was a short style guide for editing skills. When an agent makes a mistake, the natural fix was “add a sentence,” and that’s how the files got that way. Also, long inline bash blocks moved into scripts with their own tests, and extracting them turned up a couple of bugs that had been in the skill file's prose.</p>
]]></description><pubDate>Sun, 11 Oct 2026 17:46:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=50045800</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50045800</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50045800</guid></item><item><title><![CDATA[New comment by rafiss in "Five months treating bugs like patients and coding agents like a medical team"]]></title><description><![CDATA[
<p>This is a good idea; so far we've only had a few informal chats around adding something like this. Right now we play the role of Insurance Rep ourselves by manually checking on our cost dashboards and investigating expensive treatments. (As did our Director and VP on the day when we experimented with using `fable-5` agents for everything.)</p>
]]></description><pubDate>Sun, 11 Oct 2026 16:30:51 +0000</pubDate><link>https://news.ycombinator.com/item?id=50044977</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50044977</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50044977</guid></item><item><title><![CDATA[Five months treating bugs like patients and coding agents like a medical team]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.cockroachlabs.com/blog/experiment-running-hospital-code/">https://www.cockroachlabs.com/blog/experiment-running-hospital-code/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=50021899">https://news.ycombinator.com/item?id=50021899</a></p>
<p>Points: 213</p>
<p># Comments: 106</p>
]]></description><pubDate>Fri, 09 Oct 2026 15:26:33 +0000</pubDate><link>https://www.cockroachlabs.com/blog/experiment-running-hospital-code/</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=50021899</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=50021899</guid></item><item><title><![CDATA[New comment by rafiss in "A database for 2022"]]></title><description><![CDATA[
<p>> The Licensor may make an Additional Use Grant, above, permitting limited production use.<p>See the text of the Additional Use Grant for CockroachDB [1]:<p>> Additional Use Grant: You may make use of the Licensed Work, provided that you may not use the Licensed Work for a Database Service.<p>> A “Database Service” is a commercial offering that allows third parties (other than your employees an contractors) to access the functionality of the Licensed Work by creating tables whose schemas are controlled by such third parties.<p>(disclaimer: I work at Cockroach Labs)<p>[1] - <a href="https://github.com/cockroachdb/cockroach/blob/master/licenses/BSL.txt" rel="nofollow">https://github.com/cockroachdb/cockroach/blob/master/license...</a></p>
]]></description><pubDate>Mon, 04 Apr 2022 17:15:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=30909240</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=30909240</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=30909240</guid></item><item><title><![CDATA[New comment by rafiss in "Why CockroachDB and PostgreSQL Are Compatible"]]></title><description><![CDATA[
<p>We do have compatibility gaps -- some big, some small. But I would still call it compatible because it's definitely close enough to use a PostgreSQL driver with it in production. I would be curious to hear your opinion on which incompatibilities are most important to address.</p>
]]></description><pubDate>Thu, 17 Dec 2020 19:02:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=25458987</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=25458987</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=25458987</guid></item><item><title><![CDATA[New comment by rafiss in "Running Postgres in Kubernetes [pdf]"]]></title><description><![CDATA[
<p>There certainly are pain points. I don't work on this myself, but one of our other engineers wrote this blog post [0] that discusses the experience of running CockroachDB in k8s and why we chose to use it for our hosted cloud product. Another complication mentioned in there is about how to deal with the multi-region case.<p>[0] <a href="https://www.cockroachlabs.com/blog/managed-cockroachdb-on-kubernetes/" rel="nofollow">https://www.cockroachlabs.com/blog/managed-cockroachdb-on-ku...</a></p>
]]></description><pubDate>Mon, 29 Jun 2020 21:47:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=23683585</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=23683585</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=23683585</guid></item><item><title><![CDATA[New comment by rafiss in "Migrating to CockroachDB"]]></title><description><![CDATA[
<p>Our goal is to make switching over from PostgreSQL easy. For Ecto in particular, we have a tracking issue that shows the items in our backlog that would improve compatibility: <a href="https://github.com/cockroachdb/cockroach/issues/33441" rel="nofollow">https://github.com/cockroachdb/cockroach/issues/33441</a></p>
]]></description><pubDate>Sat, 08 Feb 2020 18:45:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=22277517</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=22277517</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=22277517</guid></item><item><title><![CDATA[New comment by rafiss in "How We Built a Vectorized SQL Engine"]]></title><description><![CDATA[
<p>Thanks! I do think that a columnar on-disk representation would likely help with large analytical queries. But I would imagine they would negatively affect performance of point transactions in an OLTP workload. Note that this vectorized engine we built will only be used on a given query if our SQL planner estimates that a large number of rows will be read.</p>
]]></description><pubDate>Mon, 11 Nov 2019 21:06:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=21508907</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=21508907</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=21508907</guid></item><item><title><![CDATA[How We Built a Vectorized SQL Engine]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.cockroachlabs.com/blog/how-we-built-a-vectorized-sql-engine/#">https://www.cockroachlabs.com/blog/how-we-built-a-vectorized-sql-engine/#</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=21506588">https://news.ycombinator.com/item?id=21506588</a></p>
<p>Points: 209</p>
<p># Comments: 55</p>
]]></description><pubDate>Mon, 11 Nov 2019 17:17:18 +0000</pubDate><link>https://www.cockroachlabs.com/blog/how-we-built-a-vectorized-sql-engine/#</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=21506588</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=21506588</guid></item><item><title><![CDATA[New comment by rafiss in "Docker - the Linux container runtime"]]></title><description><![CDATA[
<p>github: rafiss<p>Thanks! This might help a lot for an idea that my friends and I are working on.</p>
]]></description><pubDate>Thu, 21 Mar 2013 00:53:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=5411976</link><dc:creator>rafiss</dc:creator><comments>https://news.ycombinator.com/item?id=5411976</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=5411976</guid></item></channel></rss>