<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: dhorthy</title><link>https://news.ycombinator.com/user?id=dhorthy</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 29 Jul 2026 13:15:35 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=dhorthy" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>so wait is the finding that most of those skills reduce pass rates against SCB? wild</p>
]]></description><pubDate>Tue, 28 Jul 2026 14:55:05 +0000</pubDate><link>https://news.ycombinator.com/item?id=49084895</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49084895</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49084895</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>why would that worry you?</p>
]]></description><pubDate>Tue, 28 Jul 2026 13:31:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=49083596</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49083596</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49083596</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>agree, i think the implication is that low quality code is harder to change in the future</p>
]]></description><pubDate>Tue, 28 Jul 2026 13:30:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49083594</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49083594</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49083594</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>somebody get this man a curl-pipe-bash stat</p>
]]></description><pubDate>Tue, 28 Jul 2026 04:52:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49079507</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49079507</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49079507</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>oh i really like the idea of flipping around the order of checkpoints and comparing results. Could be an interesting way to increase/decrease difficulty even<p>i will look into how easy it would be to zip up some subset of the results without leaking anything...probably doable</p>
]]></description><pubDate>Tue, 28 Jul 2026 04:19:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49079303</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49079303</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49079303</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>Yeah I would hold that models don’t know how to simplify because most rl/benchmarks doesn’t penalize complexity</p>
]]></description><pubDate>Tue, 28 Jul 2026 03:49:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49079102</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49079102</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49079102</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>Yeah the main reason I skipped fable was because we have a ZDR with anthropic and I didn’t feel like spinning up another account to circumvent that. Next run will have fable and sol</p>
]]></description><pubDate>Tue, 28 Jul 2026 03:41:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49079051</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49079051</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49079051</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>my issue with frontier code is that it uses a model judge for quality whereas slop code bench forces a model to grapple with its own garbage code in order to receive a functionality reward</p>
]]></description><pubDate>Tue, 28 Jul 2026 02:26:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49078496</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49078496</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49078496</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>i laughed at the pelican bit its good<p>yes the labs will always prioritize the vibeslop dopamine casino as far as I can tell - making the models useful and addictive for unsophisticated users, sometimes at the expense or at the very least at the ignorance of the needs of power users</p>
]]></description><pubDate>Tue, 28 Jul 2026 02:17:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49078447</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49078447</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49078447</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>yeah someone will have to re-run this bench on various effort levels. unfortunately it is not cheap</p>
]]></description><pubDate>Tue, 28 Jul 2026 01:40:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=49078197</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49078197</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49078197</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>I agree this is an option, and the next thing on my radar is to try with a more realistic "factory-shaped" harness where you have feedback from linters and other models after each coding episode that refines the architecture.<p>For readability specifically, I've found it hard to get the models to do this with prompting. If you've talked to opus/fable for a long time on prose writing you probably felt this too</p>
]]></description><pubDate>Tue, 28 Jul 2026 01:16:56 +0000</pubDate><link>https://news.ycombinator.com/item?id=49078037</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49078037</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49078037</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>> - what "maintainable" is is probably some high dimensional space described by these signals; it'd probably require some human labeling to figure out where this space is<p>this is a nicely succinct way to put this - a multi-dimensional space where no single metric is really useful<p>state space of the system is interesting too. I would guess that for any production software with dependencies like databases/third parties that might be too hard to measure, but if you can silo off parts of your system into bounded state machines, it may be a value metric on some module behind a clean interface.<p>I think the kubernetes control loop model is a great instance of this, a handful of scoped components that own a control loop across a well-defined state machine, that can operate / recover in the face of most network partitions or downtime - the promise of CRDTs but rather more a pragmatic approach to it</p>
]]></description><pubDate>Tue, 28 Jul 2026 00:55:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=49077863</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49077863</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49077863</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>yes sol is still my daily driver for most coding tasks<p>I did find opus 5 quite handy for general knowledge work and visual design, without the cost of fable (e.g. the graphics in this post are made by opus 5)<p>but its not noticeably better than opus 4.8 in those regards, and I would not miss it if forced to go back to 4.8</p>
]]></description><pubDate>Tue, 28 Jul 2026 00:19:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49077582</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49077582</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49077582</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>no i'm spinning those up at some point this week. here's the first few prompts I used (claude opus 5 as the research orchestrator), (these were interspersed with lots of tools and assistant messages but it should get you kicked off.<p>> fetch this article for slopcodebench and help me run an eval on a subset of problems with opus 5 <a href="https://arxiv.org/html/2603.24755v1" rel="nofollow">https://arxiv.org/html/2603.24755v1</a>
> Get all the context, fetch any mentioned repos, and then propose a plan to me.<p>> i have an anthropic API key in ....
> Let's do the three challenges with Opus 4.8 and Opus 5 and Fable please. I like your minimal set. Let's try it. What do you need from me?<p>> Actually I changed my mind. I want to do two of the easy ones you picked and then I want you to pick the one with more checkpoints, maybe one of the harder ones, not the very hardest one but one with a higher number of checkpoints.<p>> Actually let's do one easy, one medium, and one hard problem please. If we have a hard problem I'd like to see that.<p>> lets rock - i think lets just do opus 4.8 and sonnet 5 and opus 5 since we our ZDR will block fable</p>
]]></description><pubDate>Tue, 28 Jul 2026 00:03:55 +0000</pubDate><link>https://news.ycombinator.com/item?id=49077412</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49077412</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49077412</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>i hope that is because you hate slop and not because you write it</p>
]]></description><pubDate>Mon, 27 Jul 2026 23:45:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49077190</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49077190</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49077190</guid></item><item><title><![CDATA[New comment by dhorthy in "Benchmarking Opus 5 on SlopCodeBench"]]></title><description><![CDATA[
<p>yeah this was just a start - the fastest cheapest thing we could try for a brand new model.<p>I'm hoping to do some more work with sol/fable in the mix as well as exploring more languages and curating the problem set to include more of the benchmark<p>I also kinda felt like opus4.5 was dumber than 4.1 personally, maybe a little biased since 4.5 was 2.5x faster and 2.5x cheaper seems to indicate its a smaller model</p>
]]></description><pubDate>Mon, 27 Jul 2026 23:42:50 +0000</pubDate><link>https://news.ycombinator.com/item?id=49077161</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49077161</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49077161</guid></item><item><title><![CDATA[Benchmarking Opus 5 on SlopCodeBench]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-opus-5-on-slop-code-bench.md">https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-opus-5-on-slop-code-bench.md</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49076391">https://news.ycombinator.com/item?id=49076391</a></p>
<p>Points: 397</p>
<p># Comments: 117</p>
]]></description><pubDate>Mon, 27 Jul 2026 22:37:52 +0000</pubDate><link>https://github.com/humanlayer/advanced-context-engineering-for-coding-agents/blob/main/benchmarking-opus-5-on-slop-code-bench.md</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49076391</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49076391</guid></item><item><title><![CDATA[New comment by dhorthy in "If coding has been solved, why does software keep getting worse?"]]></title><description><![CDATA[
<p>the one good thing about the current ios version is that it is the best it will ever be from this point forward. all future versions will be worse</p>
]]></description><pubDate>Sat, 25 Jul 2026 02:53:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49044068</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49044068</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49044068</guid></item><item><title><![CDATA[New comment by dhorthy in "Why Software Factories Fail (or: harness engineering is not enough)"]]></title><description><![CDATA[
<p>I agree this rocks I do this almost daily</p>
]]></description><pubDate>Fri, 24 Jul 2026 05:05:58 +0000</pubDate><link>https://news.ycombinator.com/item?id=49031395</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49031395</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49031395</guid></item><item><title><![CDATA[New comment by dhorthy in "Why Software Factories Fail (or: harness engineering is not enough)"]]></title><description><![CDATA[
<p>yes well said</p>
]]></description><pubDate>Fri, 24 Jul 2026 04:05:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=49031112</link><dc:creator>dhorthy</dc:creator><comments>https://news.ycombinator.com/item?id=49031112</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49031112</guid></item></channel></rss>