<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: GustavHartz</title><link>https://news.ycombinator.com/user?id=GustavHartz</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 29 Sep 2026 16:56:11 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=GustavHartz" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[Tell HN: OpenAI $500 ProMax plan listed in API]]></title><description><![CDATA[
<p>OpenAI's API is showing the first signs in the API of a new 500$ ProMax plan<p><a href="https://chatgpt.com/backend-anon/checkout_pricing_config/configs/US" rel="nofollow">https://chatgpt.com/backend-anon/checkout_pricing_config/con...</a></p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49841605">https://news.ycombinator.com/item?id=49841605</a></p>
<p>Points: 29</p>
<p># Comments: 51</p>
]]></description><pubDate>Fri, 25 Sep 2026 08:12:24 +0000</pubDate><link>https://news.ycombinator.com/item?id=49841605</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=49841605</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49841605</guid></item><item><title><![CDATA[New comment by GustavHartz in "Pi Security – Codex Security without all the bloat"]]></title><description><![CDATA[
<p>OpenAI recently released Codex-Security. We simplified it based on what we use in our own products and put it in the PI harness. We will release data on DeepSeek v4 and GLM-5.3 performance soon, but so far it's probably the best open-source vulnerability detection tool out there right now.<p>The idea is the same as Codex Security.<p>1. Build a threat model
2. Launch a lot of probing agents looking into the security based on the threat model
3. Deduplication
4. Validation
5. Severity and likelihood calibration<p>We did a deep dive article here on it as well
<a href="https://x.com/GustavHartz/status/2084926544800035266" rel="nofollow">https://x.com/GustavHartz/status/2084926544800035266</a></p>
]]></description><pubDate>Fri, 14 Aug 2026 13:52:47 +0000</pubDate><link>https://news.ycombinator.com/item?id=49298706</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=49298706</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49298706</guid></item><item><title><![CDATA[Pi Security – Codex Security without all the bloat]]></title><description><![CDATA[
<p>Article URL: <a href="https://github.com/Cecuro/open-security">https://github.com/Cecuro/open-security</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49298705">https://news.ycombinator.com/item?id=49298705</a></p>
<p>Points: 4</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 14 Aug 2026 13:52:47 +0000</pubDate><link>https://github.com/Cecuro/open-security</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=49298705</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49298705</guid></item><item><title><![CDATA[Under the Hood of Codex Security]]></title><description><![CDATA[
<p>Article URL: <a href="https://twitter.com/GustavHartz/status/2084926544800035266">https://twitter.com/GustavHartz/status/2084926544800035266</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=49180335">https://news.ycombinator.com/item?id=49180335</a></p>
<p>Points: 5</p>
<p># Comments: 2</p>
]]></description><pubDate>Wed, 05 Aug 2026 09:04:12 +0000</pubDate><link>https://twitter.com/GustavHartz/status/2084926544800035266</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=49180335</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49180335</guid></item><item><title><![CDATA[New comment by GustavHartz in "Codex Security"]]></title><description><![CDATA[
<p>TLDR on how it works: It's a small stack of skill files and some JS code that starts a large number of Codex sessions. They all get the prompt and the same scope with a limited set of tools. Not one prompt in the repo contains security advice, performance is obtained only through scaling the number of agents looking at the code</p>
]]></description><pubDate>Wed, 05 Aug 2026 09:02:38 +0000</pubDate><link>https://news.ycombinator.com/item?id=49180328</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=49180328</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49180328</guid></item><item><title><![CDATA[What 650k commits say about how crypto bugs change]]></title><description><![CDATA[
<p>Article URL: <a href="https://twitter.com/GustavHartz/status/2073819470796128402">https://twitter.com/GustavHartz/status/2073819470796128402</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48802836">https://news.ycombinator.com/item?id=48802836</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Mon, 06 Jul 2026 10:30:58 +0000</pubDate><link>https://twitter.com/GustavHartz/status/2073819470796128402</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=48802836</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48802836</guid></item><item><title><![CDATA[More agents are better than fine-tuned model for pen-testing]]></title><description><![CDATA[
<p>Article URL: <a href="https://twitter.com/GustavHartz/status/2072294954404135275">https://twitter.com/GustavHartz/status/2072294954404135275</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48745599">https://news.ycombinator.com/item?id=48745599</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Wed, 01 Jul 2026 12:25:59 +0000</pubDate><link>https://twitter.com/GustavHartz/status/2072294954404135275</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=48745599</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48745599</guid></item><item><title><![CDATA[New comment by GustavHartz in "Building effective pen-testing agents"]]></title><description><![CDATA[
<p>This started as a response to the recent "you have to post-train a model to pen-test" Show HN — we don't think you need to, just makes life a bit easier.<p>Across 10K+ of our agent transcripts from benchmarking against OpenAI's EVMBench, we saw zero refusals. In the closed-frontier models, the refusal you hit is mostly a separate content classifier, or a system prompt, not so much the model itself. Breadth (more cheap agents) beats a bigger model, but it puts more requirements on context engineering<p><a href="https://news.ycombinator.com/item?id=48609231">https://news.ycombinator.com/item?id=48609231</a></p>
]]></description><pubDate>Fri, 26 Jun 2026 08:16:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48683860</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=48683860</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48683860</guid></item><item><title><![CDATA[Building effective pen-testing agents]]></title><description><![CDATA[
<p>Article URL: <a href="https://cecuro.ai/blog/building-effective-pen-testing-agents">https://cecuro.ai/blog/building-effective-pen-testing-agents</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=48683859">https://news.ycombinator.com/item?id=48683859</a></p>
<p>Points: 5</p>
<p># Comments: 1</p>
]]></description><pubDate>Fri, 26 Jun 2026 08:16:30 +0000</pubDate><link>https://cecuro.ai/blog/building-effective-pen-testing-agents</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=48683859</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48683859</guid></item><item><title><![CDATA[New comment by GustavHartz in "In 92% of DeFi exploits AI security review flags underlying problem"]]></title><description><![CDATA[
<p>Performance has gotten a lot better the last 6 months, at a level where we almost don't see it anymore at Cecuro.ai. PoC generation and multiple validation agents debating validity is the key differentiator. This is an ok paper on the topic <a href="https://arxiv.org/abs/2511.02780" rel="nofollow">https://arxiv.org/abs/2511.02780</a></p>
]]></description><pubDate>Mon, 23 Feb 2026 12:17:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=47121332</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=47121332</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47121332</guid></item><item><title><![CDATA[In 92% of DeFi exploits AI security review flags underlying problem]]></title><description><![CDATA[
<p>Article URL: <a href="https://www.coindesk.com/business/2026/02/20/specialized-ai-detects-92-of-real-world-defi-exploits">https://www.coindesk.com/business/2026/02/20/specialized-ai-detects-92-of-real-world-defi-exploits</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47110937">https://news.ycombinator.com/item?id=47110937</a></p>
<p>Points: 3</p>
<p># Comments: 2</p>
]]></description><pubDate>Sun, 22 Feb 2026 13:43:37 +0000</pubDate><link>https://www.coindesk.com/business/2026/02/20/specialized-ai-detects-92-of-real-world-defi-exploits</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=47110937</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47110937</guid></item><item><title><![CDATA[OAI: EVM Bench LLM Accuracy on Smart Contract Review and Pentesting]]></title><description><![CDATA[
<p>Article URL: <a href="https://openai.com/index/introducing-evmbench/">https://openai.com/index/introducing-evmbench/</a></p>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=47071489">https://news.ycombinator.com/item?id=47071489</a></p>
<p>Points: 2</p>
<p># Comments: 0</p>
]]></description><pubDate>Thu, 19 Feb 2026 08:47:51 +0000</pubDate><link>https://openai.com/index/introducing-evmbench/</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=47071489</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=47071489</guid></item><item><title><![CDATA[New comment by GustavHartz in "AI agents find $4.6M in blockchain smart contract exploits"]]></title><description><![CDATA[
<p>We've been working on this at cecuro.ai. When we test Sonnet 4.5 against real cyber security audit reports from the major firms on code that came out after the model was trained, it finds around 95% of the same bugs the auditors found. Also catches some medium severity stuff they missed. We find that you can't just point one model at a contract and expect good results though. Need to run multiple models with different prompts because they each have different blind spots. Still tricky to get working well and not cheap. Happy to share more if anyone's curious</p>
]]></description><pubDate>Sun, 07 Dec 2025 16:42:22 +0000</pubDate><link>https://news.ycombinator.com/item?id=46183003</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=46183003</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=46183003</guid></item><item><title><![CDATA[Ask HN: Do you think MS Copilot will replace junior consultants?]]></title><description><![CDATA[
<p>As a consultant, I spend way too much time buried in emails, teams, Slack threads, and random SharePoint folders trying to come up with answers to questions like:
 • What’s the current project status?
 • What is that XXX deck?
 • What is the revised timeline?<p>Junior consultants at my company usually get the job of building timelines, trackers, and endless update slides. I’m curious whether MS Copilot, due to the data access, could fill this role—parsing all comms, tracking deliverables, and giving you clean overviews, action items, and even draft decks.<p>Do you think Copilot will evolve into this kind of tool? Or is there still a gap here that a startup could fill? Curious if others are seeing the same pain, or already building around it.<p>Idea visualized: https://theo-consult-canvas.lovable.app/dashboard</p>
<hr>
<p>Comments URL: <a href="https://news.ycombinator.com/item?id=43673131">https://news.ycombinator.com/item?id=43673131</a></p>
<p>Points: 1</p>
<p># Comments: 0</p>
]]></description><pubDate>Sun, 13 Apr 2025 14:37:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=43673131</link><dc:creator>GustavHartz</dc:creator><comments>https://news.ycombinator.com/item?id=43673131</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=43673131</guid></item></channel></rss>