<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: jascha_eng</title><link>https://news.ycombinator.com/user?id=jascha_eng</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Sat, 05 Sep 2026 06:45:04 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=jascha_eng" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by jascha_eng in "Artificial Analysis Intelligence Index v4.2"]]></title><description><![CDATA[
<p>Imo the omniscience index they have has the highest correlation to actual usefulness of the models.<p><a href="https://artificialanalysis.ai/evaluations/omniscience" rel="nofollow">https://artificialanalysis.ai/evaluations/omniscience</a><p>> measures knowledge reliability and hallucination. It rewards correct answers, penalizes hallucinations, and has no penalty for refusing to answer.<p>This is so useful because it makes you actually trust a models output. A high score on benchmarks is not as useful because a model overtrained to always answer will give confidently wrong responses. But this index measures how often it is correct while penalizing wrong responses so that a high score means you can trust this model more and when it doesn't know it is more likely to tell you that it really doesn't know rather than making shit up.<p>Fable also performs a lot better than opus 5 here which correlates very strongly with perceived strength despite the models performing similarly on e.g. DeepSWE<p>Astra is a big jump from sol and performs the same or slightly better than fable here.</p>
]]></description><pubDate>Sat, 05 Sep 2026 01:27:41 +0000</pubDate><link>https://news.ycombinator.com/item?id=49572176</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49572176</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49572176</guid></item><item><title><![CDATA[New comment by jascha_eng in "What Is a Harness?"]]></title><description><![CDATA[
<p>The ai hype word for 2026 after agent in 2025 for any LLM powered application.<p>Well kind of, I wouldn't be surprised to see that some things marketed as agents are actually good old deterministic software.</p>
]]></description><pubDate>Sun, 23 Aug 2026 15:02:53 +0000</pubDate><link>https://news.ycombinator.com/item?id=49409388</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49409388</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49409388</guid></item><item><title><![CDATA[New comment by jascha_eng in "The August 17 outage"]]></title><description><![CDATA[
<p>Instagram, Facebook and even threads all had much more mundane growth rates and definitely no unexpected jumps like GitHub is experiencing. I'm sure if suddenly the solar system had 10 more earths with each about 10 billion people and they would all start using Instagram tomorrow we would have exactly the same growing pains and outages that GitHub has today.<p>Luckily for Meta agents are not yet as much into doomscrolling as humans are.</p>
]]></description><pubDate>Fri, 21 Aug 2026 11:42:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49386651</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49386651</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49386651</guid></item><item><title><![CDATA[New comment by jascha_eng in "Humans missed 1 in 3 threats approving AI agent commands across 40k game runs"]]></title><description><![CDATA[
<p>1 in 3 is not terrible you just need a few more humans in the loop to reduce the error rate meaningfully. Combined with other classifier models and heuristics you can get good results. Humans can probably also perform better if they don't have to judge every single command but just suspicious ones our attention is limited after all.</p>
]]></description><pubDate>Thu, 06 Aug 2026 13:45:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49196614</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49196614</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49196614</guid></item><item><title><![CDATA[New comment by jascha_eng in "How to Spot AI Writing"]]></title><description><![CDATA[
<p>But fables is often worth diving into at least. Opus 5 says this and then rambles on about something completely irrelevant or even wrong</p>
]]></description><pubDate>Sat, 01 Aug 2026 14:55:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49135003</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49135003</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49135003</guid></item><item><title><![CDATA[New comment by jascha_eng in "Writing by hand is good for your brain"]]></title><description><![CDATA[
<p>The study he cites is also specifically using digital pens<p>> Brain electrical activity was recorded in 36 university students as they were handwriting visually presented words using a digital pen and typewriting the words on a keyboard<p>kinda funny.<p>It's also not clear from this study that there is actually any benefit to "more brain activity". Of course doing more complex motor tasks requires more brain activity but nobody guarantees that this helps with remembering things.</p>
]]></description><pubDate>Thu, 23 Jul 2026 19:07:32 +0000</pubDate><link>https://news.ycombinator.com/item?id=49026569</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49026569</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49026569</guid></item><item><title><![CDATA[New comment by jascha_eng in "OverpAId – Fire your CEO. Hire the future"]]></title><description><![CDATA[
<p>tbh AI is great at vague strategic decision making with a low risk bias. I wouldn't mind working for Claude</p>
]]></description><pubDate>Wed, 22 Jul 2026 12:28:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=49005693</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=49005693</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49005693</guid></item><item><title><![CDATA[New comment by jascha_eng in "Kimi Work"]]></title><description><![CDATA[
<p>huh? because im curious what they used?
Claude Code and codex take completely different approaches. The core loop of feeding generation and having a bash tool is entirely the same. If you want to start building your own thing I'm sure there is good setups out there to start from rather than going from zero.<p>Ofc I can just fork gemini-cli if I want my own version but like I don't think that's what OP meant by build your own thing and customize it.</p>
]]></description><pubDate>Tue, 21 Jul 2026 23:21:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=48999710</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48999710</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48999710</guid></item><item><title><![CDATA[New comment by jascha_eng in "Kimi Work"]]></title><description><![CDATA[
<p>How did you build your own coding agent? What language/framework did you use?</p>
]]></description><pubDate>Mon, 20 Jul 2026 23:31:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=48986231</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48986231</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48986231</guid></item><item><title><![CDATA[New comment by jascha_eng in "Typing Speed Test, but for Developers"]]></title><description><![CDATA[
<p>It used to be that DevOps was the movement of merging Ops leftwards so that Devs own the Operation of their software.</p>
]]></description><pubDate>Sat, 18 Jul 2026 21:27:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=48962592</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48962592</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48962592</guid></item><item><title><![CDATA[New comment by jascha_eng in "Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?"]]></title><description><![CDATA[
<p>And that doesn't work with a simple prompt?</p>
]]></description><pubDate>Sat, 18 Jul 2026 16:34:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48959607</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48959607</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48959607</guid></item><item><title><![CDATA[New comment by jascha_eng in "Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?"]]></title><description><![CDATA[
<p>Can you give an example? And more curious about what you do with the resulting code afterwards I imagine its gonna be a big chunk then?</p>
]]></description><pubDate>Sat, 18 Jul 2026 14:28:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=48958475</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48958475</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48958475</guid></item><item><title><![CDATA[New comment by jascha_eng in "Fable 5 vs. GPT-5.6 Sol on an NP-Hard Problem: Does /goal help?"]]></title><description><![CDATA[
<p>Is this useful? I feel like the problem is usually not that the model isn't capable of achieving what I give it, but the way it does it. Especially if originally I didn't 100% know how I would do it myself the model often takes weird paths through the code base, takes shortcuts that end up in weird feature interactions or pulls in a dependency without weighting if it could've been done without that.<p>I haven't really found a good way to solve this other than:<p>1. Produce an initial PR fulfilling all the requirements I knew at the start<p>2. Chat with the model about any weird snippets I notice and talk through alternatives<p>3. Simplify anything that I think is overengineered or plain unncessary<p>Sometimes I restart all over with more precise requirements but then it sometimes makes different mistakes/takes different shortcuts.<p>In practice the earlier I review the better the end result imo, so /goal seems very unproductive to me?</p>
]]></description><pubDate>Sat, 18 Jul 2026 13:55:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=48958178</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48958178</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48958178</guid></item><item><title><![CDATA[New comment by jascha_eng in "Claude Code: Anatomy of a Misfeature"]]></title><description><![CDATA[
<p>Until the classifier is wrong or also prompt injected. the classifier is just as vulnerable as the model itself is. Yes it is harder to break but trying to make a nondeterministic tool deterministic by adding another nondeterministic one on top just reduces the chance of something going wrong.<p>Tbf as long as that chance is low enough it doesn't matter in practice, but I have definitely seen the classifier approve things that were questionable, and I've also seen it decline things that were obviously okay.</p>
]]></description><pubDate>Fri, 17 Jul 2026 17:06:07 +0000</pubDate><link>https://news.ycombinator.com/item?id=48949656</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48949656</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48949656</guid></item><item><title><![CDATA[New comment by jascha_eng in "OnePlus halts operations in USA and Europe"]]></title><description><![CDATA[
<p>But a pixel is quite a bit more expensive no? At that point you can consider an iPhone?</p>
]]></description><pubDate>Thu, 16 Jul 2026 13:52:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48934604</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48934604</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48934604</guid></item><item><title><![CDATA[New comment by jascha_eng in "OnePlus halts operations in USA and Europe"]]></title><description><![CDATA[
<p>Sad I have a 6 year old oneplus and was looking for a new phone somewhat soon, would've considered them again for sure. Any alternatives? They always had a reputation for me for being a great no fuss, little bloat and simply fast android phone.</p>
]]></description><pubDate>Thu, 16 Jul 2026 13:18:23 +0000</pubDate><link>https://news.ycombinator.com/item?id=48934134</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48934134</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48934134</guid></item><item><title><![CDATA[New comment by jascha_eng in "The real prices of frontier models"]]></title><description><![CDATA[
<p>I mean it might lead to better performance on the model side. So the tokenizer is better but more expensive.</p>
]]></description><pubDate>Mon, 13 Jul 2026 21:43:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48899315</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48899315</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48899315</guid></item><item><title><![CDATA[New comment by jascha_eng in "The real prices of frontier models"]]></title><description><![CDATA[
<p>Aside of the claudeisms and the obvious AI smell, it overexplains everything and doesn't come to any useful conclusions. It's just not a good post.<p>The nudge to think about both "tokenization as variable" as well as actual tokens consumed per task is still good.</p>
]]></description><pubDate>Mon, 13 Jul 2026 21:40:43 +0000</pubDate><link>https://news.ycombinator.com/item?id=48899288</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48899288</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48899288</guid></item><item><title><![CDATA[New comment by jascha_eng in "Ask HN: What Are You Working On? (July 2026)"]]></title><description><![CDATA[
<p>Slowly improving the UX on my SQL review/approval tool: <a href="https://github.com/kviklet/kviklet" rel="nofollow">https://github.com/kviklet/kviklet</a><p>Also finally closed the first real customer on it recently!<p>I want to get through a large chunk of the open issues the next few weeks and then spend some time building agentic capabilities for it.
I believe a central place to configure database access for your dev team without having to share passwords and with sensible review policies should also help e.g. if claude needs to access production data to validate a premise.<p>Still have to figure out the right UX though not sure the agent should have the exact same review requirements that a human does. Maybe it needs to be configurable separately</p>
]]></description><pubDate>Mon, 13 Jul 2026 02:25:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=48887115</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48887115</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48887115</guid></item><item><title><![CDATA[New comment by jascha_eng in "GPT-5.6 Sol Ultra will be in Codex"]]></title><description><![CDATA[
<p>im talking about anthropics pricing</p>
]]></description><pubDate>Tue, 07 Jul 2026 18:00:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=48821283</link><dc:creator>jascha_eng</dc:creator><comments>https://news.ycombinator.com/item?id=48821283</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48821283</guid></item></channel></rss>