<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: screm</title><link>https://news.ycombinator.com/user?id=screm</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Fri, 04 Sep 2026 06:22:55 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=screm" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>It should be better now, including in mobile, thanks both for the feedback!</p>
]]></description><pubDate>Thu, 03 Sep 2026 23:42:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558650</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558650</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558650</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Yep makes sense I’m relaxing them</p>
]]></description><pubDate>Thu, 03 Sep 2026 23:09:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558383</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558383</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558383</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Haha there is a lot at stake for sure</p>
]]></description><pubDate>Thu, 03 Sep 2026 23:05:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558356</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558356</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558356</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Yeah sounds kind of like the equivalent of SEA for AI agents (AEA?) except that it’s sneakier since agents can act without you noticing.. anyway this is in the hands of the labs</p>
]]></description><pubDate>Thu, 03 Sep 2026 23:04:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558345</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558345</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558345</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Exactly!</p>
]]></description><pubDate>Thu, 03 Sep 2026 22:38:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558096</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558096</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558096</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Hey, thanks for the feedback, the leaderboards aren't displaying well on mobile indeed, we are currently shipping a fix that should help with that. Thanks anyway!</p>
]]></description><pubDate>Thu, 03 Sep 2026 22:38:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558090</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558090</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558090</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Definitely! But about concentration I'm not so sure, there are ways to counter this effect so in the end it will be a fight like SEO is today. What is certain though is that getting recommended by coding agents will be a top prio for all dev tools.</p>
]]></description><pubDate>Thu, 03 Sep 2026 22:35:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558075</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558075</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558075</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Not sure I got your question right but if you are wondering for your own coding agent then I guess the answer would be a skill?
Here what I meant by "how to influence coding agents choices and get products picked" is from a vendor PoV, making sure any developer x codebase in the world asking for a tool in your category gets your tool recommended and implemented by the coding agent.</p>
]]></description><pubDate>Thu, 03 Sep 2026 22:33:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49558054</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49558054</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49558054</guid></item><item><title><![CDATA[New comment by screm in "Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out"]]></title><description><![CDATA[
<p>Hey!<p>Disclaimer: I am a Co-Founder of Armature (YC P26) which sells growth services to dev tools. This study is part of our broader work on how to influence coding agents choices and get products picked.<p>To understand how agents pick tools we measured close to 17k sessions on an environment where agents run exactly like in the real world, on various repositories, talking to different personas (vibe-coder, junior or senior engineers) in different sizes of companies.<p>All the results are now public and we'd love to know what findings surprise you the most, here are a few we found interesting:
- Claude Code rarely searches the web while Codex almost always does it and Cursor sits in the middle.
- Coding agents disagree more frequently than they agree.
- Some players (LangChain, Supabase, Netlify, Paypal, Adyen) are almost always mentioned in their categories but never chosen.
- Modifying repository context can change the pick entirely.<p>If you feel like digging, all the traces are there and we probably missed interesting learnings so let us know what you find!</p>
]]></description><pubDate>Thu, 03 Sep 2026 21:22:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49557233</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49557233</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49557233</guid></item><item><title><![CDATA[Which tools do Claude, Codex and Cursor choose? We measured 17k runs to find out]]></title><description><![CDATA[
<p>Article URL: <a href="https://armature.tech/blog/which-tools-coding-agents-install">https://armature.tech/blog/which-tools-coding-agents-install</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49557206">https://news.ycombinator.com/item?id=49557206</a></p>
<p>Points: 166</p>
<p># Comments: 58</p>
]]></description><pubDate>Thu, 03 Sep 2026 21:20:34 +0000</pubDate><link>https://armature.tech/blog/which-tools-coding-agents-install</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49557206</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49557206</guid></item><item><title><![CDATA[New comment by screm in "Show HN: Product analytics (and evals) for agent sessions on your MCP"]]></title><description><![CDATA[
<p>Makes sense, just feel free to reach out if you think any of our features could become useful at some point or just want to discuss "building for agents"!<p>Btw we recently shipped [evals](<a href="https://armature.tech/blog/armature-launch-evals-for-mcps-and-clis">https://armature.tech/blog/armature-launch-evals-for-mcps-an...</a>) and are considering supporting local MCPs too so let me know if we should!</p>
]]></description><pubDate>Thu, 06 Aug 2026 15:25:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=49197997</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49197997</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49197997</guid></item><item><title><![CDATA[New comment by screm in "Show HN: Product analytics (and evals) for agent sessions on your MCP"]]></title><description><![CDATA[
<p>Yes it does because we can't see the real chat transcripts or any kind of history or memory. We only make sure the calls to your MCP server include "brief, task-specific user intent" following OpenAI's Apps SDK guidelines here: <a href="https://developers.openai.com/plugins/app-guidelines" rel="nofollow">https://developers.openai.com/plugins/app-guidelines</a></p>
]]></description><pubDate>Thu, 06 Aug 2026 08:32:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49194065</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49194065</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49194065</guid></item><item><title><![CDATA[New comment by screm in "Show HN: Product analytics (and evals) for agent sessions on your MCP"]]></title><description><![CDATA[
<p>Thanks! Happy to give you a tour and see if it can be helpful or just discuss how you handle these challenges on your side!</p>
]]></description><pubDate>Thu, 06 Aug 2026 08:13:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49193933</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49193933</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49193933</guid></item><item><title><![CDATA[New comment by screm in "Show HN: Product analytics (and evals) for agent sessions on your MCP"]]></title><description><![CDATA[
<p>Thanks, great questions!<p>1. We ask it :) The SDK adds an optional "telemetry" object to each tool's input schema, with fields like user_intent, agent_thinking and user_frustration. Then the calling agent just fills them in as part of the tool call, and the SDK strips the block before your handler runs, so your business logic never sees it!<p>2. Yes! for example if your server uses the official MCP SDK it is literally:<p><pre><code>  import { createMcpAnalyticsServer } from "@armature-tech/mcp-analytics";
  import { createMyMcpServer } from "./server.js";

  const server = createMcpAnalyticsServer(createMyMcpServer);
</code></pre>
-> We explain everything in <a href="https://docs.armature.tech">https://docs.armature.tech</a> but let me know if anything's unclear!<p>3. Yes the SDKs are open source in TypeScript, Python and Go. What we mean by "client side" is that the SDK runs inside your MCP server process so nothing runs on your end users' devices, their client just sees one extra optional field in your tool schemas. The only things that get sent are: tool name, timing, outcome, session and actor identifiers, client user agent, the telemetry fields the agent chose to send, and truncated previews of inputs and outputs after sanitization. And yes, all this is configurable: redact lets you plug your own redaction into previews, redactEvent can rewrite or drop whole events, captureTelemetry: false disables all conversation-derived data, and enabled: false turns the whole thing off!<p>4. A simple way to see it is to think OTel vs PostHog on a regular web app. OTel is on the observability side (what your server did, traces, latency, errors) while PostHog is on the product analytics side (what users are trying to do, whether they succeed, where they drop off). We can't replace PostHog with API call logs so it's the same for MCPs. So Armature bridges this gap on the product analytics side. What makes it harder with MCPs is that the product half (the user's prompt, the agent's reasoning and the frustration) is not in your logs because it lives in your users' AI client not on your server. So we need to reconstruct these sessions and then on top of it we can build the analytics layer and add evals.
About BrainFuse, do you mean Langfuse / Braintrust LLM observability tools? If yes, those instrument an agent you own and run while with an MCP you are on the other side: someone else's agent is calling you, and you cannot instrument their side with these tools :/<p>Hope it answers and thanks for the kind words!</p>
]]></description><pubDate>Mon, 03 Aug 2026 20:26:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49160967</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49160967</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49160967</guid></item><item><title><![CDATA[Show HN: Product analytics (and evals) for agent sessions on your MCP]]></title><description><![CDATA[
<p>Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent to do and what the agent thought.<p>You wrap your MCP in 3 lines of code (our SDK is available in Typescript, Python and Go) and start seeing in your dashboard:
- All sessions reconstructed: it’s like reading the real conversation the user had inside Claude or ChatGPT!
- A ranking of your MCP most popular use cases, built from sessions clustering
- The most frequent issues your users’ agents encounter so you can fix them.<p>Here is a quick demo: <a href="https://youtu.be/ZFlvquhyNMQ" rel="nofollow">https://youtu.be/ZFlvquhyNMQ</a><p>The story behind this is that we initially launched Armature as a standalone testing tool (<a href="https://www.ycombinator.com/launches/QQc-armature-making-your-app-finally-usable-by-ai-agents">https://www.ycombinator.com/launches/QQc-armature-making-you...</a>) that could naturally be used through an MCP itself. We quickly realized we had no idea how our users were using Armature MCP and if they were satisfied with it or frustrated. It’s something we had also experienced in our previous companies: Louis built MCPs exposed to millions of users and Theo was a Forward Deployed Engineer at Palantir before joining a Datadog spin-off as Founding Engineer. Both testing and product analytics had always been real pains when exposing a product to agents but we always thought there wasn’t much we could do about analytics because the conversation lived in our users’ AI client.<p>Then it struck us: what if we asked the agents why they were making this or that tool call? And what’s the user's intent or potential frustration? So we started experimenting with MCP instrumentation and the use-cases actually surprised us! Many of our first customers had implemented workarounds for their CI to trigger new tests or for their coding agents to fetch the results efficiently. Even though we talked to our first users regularly, they had never shared this feedback with us. We then built automations to automatically cluster use-cases, identify issues frequently encountered and let our own coding agents fix them. When our CTO friends heard about this, they wanted to try it for themselves so we gave them access to a cloned version of our internal product and they started sharing feedback like they never did on our “real” product!<p>That’s when we decided to start working seriously on MCP Analytics as a product. At first we were afraid of degrading MCP performance so we iterated until we reached the exact same success rate as without our instrumentation (89.17 % vs 89.15 % pass rate out of 870 runs). Then privacy was an obvious constraint so we applied the same methods we had learned from working with banking data or building sensitive data scanning in logs. Today, redaction runs client-side before reaching our servers. There are still a lot of things we haven’t fully figured out: not all fields are equally filled by all models, session fingerprinting for serverless / stateless MCPs isn’t perfect, and use-case clustering remains to be optimized.<p>But we are finally launching our analytics product to everyone, self-serve at <a href="https://armature.tech">https://armature.tech</a> with a set-up that takes less than 5 minutes and a generous free tier.<p>And now we are working on fully closing the loop, bringing evals back in our product so we can: identify top workflows and issues -> recommend fixes and improvements -> test fixes at scale on the same workflows run by users, across all harnesses and models -> open PRs to ship fixes directly.
The evals can be generated automatically from the session analytics so you can catch every regression and can test every improvement’s real impact across all models and harnesses before shipping it.<p>Here’s an example to make it more concrete: 10 days ago, a marketing automation platform which has had early access to what we built for weeks identified thanks to MCP Analytics that users were frustrated not being able to change their target audience after campaign creation. So they shipped the feature and tested it successfully locally with Claude Code on Fable 5. Then a few days later when preparing their new MCP public release, they ran a suite of evals on Armature and realized that small models could hallucinate audience_ids which would lead their MCP to send the campaign to ALL their contacts by default (which could obviously lead to disasters in prod). This is the kind of story that makes what we are building feel so helpful!<p>Now, the most useful feedback for us would be to know what’s still missing in our product so you can feel you are now in full control of the “Agent Experience”.
And if you run an MCP in production we’d also love to know: what do you do today to know if agents succeed and if the users behind them are happy?</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49157807">https://news.ycombinator.com/item?id=49157807</a></p>
<p>Points: 42</p>
<p># Comments: 8</p>
]]></description><pubDate>Mon, 03 Aug 2026 16:17:59 +0000</pubDate><link>https://armature.tech/</link><dc:creator>screm</dc:creator><comments>https://news.ycombinator.com/item?id=49157807</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49157807</guid></item></channel></rss>