<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: itintheory</title><link>https://news.ycombinator.com/user?id=itintheory</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Wed, 07 Oct 2026 01:56:33 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=itintheory" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by itintheory in "Example.com just launched the biggest redesign in decades"]]></title><description><![CDATA[
<p>It's static content in Cloudflare CDN.  Seems unlikely that many requests actually hit any kind of origin server. If I were building something like this the origin would be Cloudflare R2 (or similar) object storage, with a very long cache lifetime.  Nothing to update, nothing to maintain.<p>But beyond that, what are you referring to as compression?  When I look at the source I see a VERY short easily human-readable HTML page.  It doesn't have newlines, but I wouldn't consider that compression.  It also references a short, but non-obfuscated javascript file at <a href="https://example.com/s.js" rel="nofollow">https://example.com/s.js</a> .<p>Maybe I misunderstood what you're referring to...</p>
]]></description><pubDate>Tue, 06 Oct 2026 13:05:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49977822</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49977822</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49977822</guid></item><item><title><![CDATA[New comment by itintheory in "Example.com just launched the biggest redesign in decades"]]></title><description><![CDATA[
<p>> standard windows way<p>The thread you linked to suggests 'timeout'[0]<p>[0] <a href="https://learn.microsoft.com/en-us/windows-server/administration/windows-commands/timeout" rel="nofollow">https://learn.microsoft.com/en-us/windows-server/administrat...</a></p>
]]></description><pubDate>Tue, 06 Oct 2026 12:58:18 +0000</pubDate><link>https://news.ycombinator.com/item?id=49977740</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49977740</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49977740</guid></item><item><title><![CDATA[New comment by itintheory in "An agent used DNS to reach an external chatbot"]]></title><description><![CDATA[
<p>But that's still DNS, right?  Where does it bleed over into an LLM API?  I understand there are DNS to LLM server projects, but how would the agent discover one? And I'm guessing most people who run something like that don't expose it publicly...</p>
]]></description><pubDate>Sun, 27 Sep 2026 01:17:20 +0000</pubDate><link>https://news.ycombinator.com/item?id=49862360</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49862360</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49862360</guid></item><item><title><![CDATA[New comment by itintheory in "An agent used DNS to reach an external chatbot"]]></title><description><![CDATA[
<p>Right, but your have to run this server somewhere, which the agent couldn't do.</p>
]]></description><pubDate>Sun, 27 Sep 2026 01:14:11 +0000</pubDate><link>https://news.ycombinator.com/item?id=49862337</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49862337</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49862337</guid></item><item><title><![CDATA[New comment by itintheory in "An agent used DNS to reach an external chatbot"]]></title><description><![CDATA[
<p>What DNS service did the agent discover that allowed it to execute arbitrary llm queries?  And how?</p>
]]></description><pubDate>Sun, 27 Sep 2026 00:08:21 +0000</pubDate><link>https://news.ycombinator.com/item?id=49861886</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49861886</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49861886</guid></item><item><title><![CDATA[New comment by itintheory in "An agent used DNS to reach an external chatbot"]]></title><description><![CDATA[
<p>Seems like this requires operating a proxy somewhere.  In TFA it seems like all they needed was a DNS client, but I'm not at all clear how that could work.  I'm definitely curious about the technique though.</p>
]]></description><pubDate>Sun, 27 Sep 2026 00:01:37 +0000</pubDate><link>https://news.ycombinator.com/item?id=49861842</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49861842</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49861842</guid></item><item><title><![CDATA[New comment by itintheory in "Plan mode is dead"]]></title><description><![CDATA[
<p>On the quick start page it looks like you just pass it as a CLI argument: <a href="https://docs.plannotator.ai/open-source/start/quickstart" rel="nofollow">https://docs.plannotator.ai/open-source/start/quickstart</a></p>
]]></description><pubDate>Sat, 26 Sep 2026 14:28:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49856936</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49856936</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49856936</guid></item><item><title><![CDATA[New comment by itintheory in "Plan mode is dead"]]></title><description><![CDATA[
<p>Couldn't you commit the plan markdown to git to get the tracking you're referring to?</p>
]]></description><pubDate>Sat, 26 Sep 2026 14:27:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=49856914</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49856914</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49856914</guid></item><item><title><![CDATA[New comment by itintheory in "CrowdSec Source Code Leak"]]></title><description><![CDATA[
<p>What kind of fingerprinting are you thinking of?  JA4?  I haven't found a way to do that inexpensively at our scale, but we may have to go that route - looking at CloudFront bot mitigation.<p>For behavioral, we have Anubis honeypot functionality turned on, but it doesn't seem to be effective for 99% of scrapers.  Anubis is also running behind TLS termination, so I don't think it can do full JA4.  It does have the less robust JA4H apparently, but I'm not sure how effective that will be.<p>Edit: Oh yeah, forgot to mention - it's almost 100% residential proxies.  Primarily China Telecom and China Unicom.  Unfortunately those providers are HUGE and also host a ton of legitimate users all over Asia.</p>
]]></description><pubDate>Thu, 17 Sep 2026 20:28:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49746059</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49746059</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49746059</guid></item><item><title><![CDATA[New comment by itintheory in "CrowdSec Source Code Leak"]]></title><description><![CDATA[
<p>This was just blocklist based.  We had the main community list and a handful of the curated paid lists enabled.<p>wrt bot detection - this sounds very much like Anubis which we're also using with some success.</p>
]]></description><pubDate>Thu, 17 Sep 2026 17:53:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=49744246</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49744246</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49744246</guid></item><item><title><![CDATA[New comment by itintheory in "CrowdSec Source Code Leak"]]></title><description><![CDATA[
<p>We had the main community blocklist and several of their pricey paid blocklists enabled in a PoC capacity.  We had a lot of legitimate users end up blocked.  In some cases these may have been VPN exit nodes, or users on CG-NAT, or devices on a shared network with some other compromised / bot device.  I didn't get 100% of the details, just that we were inundated with support requests from real users that ended up blocked.</p>
]]></description><pubDate>Thu, 17 Sep 2026 17:52:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49744241</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49744241</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49744241</guid></item><item><title><![CDATA[New comment by itintheory in "CrowdSec Source Code Leak"]]></title><description><![CDATA[
<p>The basic software is open source, and the list is free if you're running the tool and contributing detections back.  They do have some curated lists that you have to pay for.<p>It's quite a bit less expensive than most other commercial products of this kind that I've looked at.</p>
]]></description><pubDate>Thu, 17 Sep 2026 17:03:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49743598</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49743598</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49743598</guid></item><item><title><![CDATA[New comment by itintheory in "CrowdSec Source Code Leak"]]></title><description><![CDATA[
<p>We implemented CrowdSec for bot/scraping mitigation.  The architecture is sound, but it ended up having an unacceptable false positive rate for us.  This may be an issue with any kind of IP reputation approach.  After a couple of months of work getting it ready to go I had to turn it off after a couple of days.</p>
]]></description><pubDate>Thu, 17 Sep 2026 16:46:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=49743364</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49743364</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49743364</guid></item><item><title><![CDATA[New comment by itintheory in "Neovim have a ~$800k Bitcoin donation sitting untouched since 2023"]]></title><description><![CDATA[
<p>Goldbacks have entered the chat. [0]<p>[0] <a href="https://www.goldback.com/" rel="nofollow">https://www.goldback.com/</a></p>
]]></description><pubDate>Thu, 17 Sep 2026 13:31:02 +0000</pubDate><link>https://news.ycombinator.com/item?id=49740511</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49740511</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49740511</guid></item><item><title><![CDATA[New comment by itintheory in "An update on Wayback Machine access"]]></title><description><![CDATA[
<p>The data <i>IS</i> available.  You can download it all from several sources in one big dump.  And yet we're still scraped.</p>
]]></description><pubDate>Wed, 16 Sep 2026 19:10:40 +0000</pubDate><link>https://news.ycombinator.com/item?id=49731531</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49731531</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49731531</guid></item><item><title><![CDATA[New comment by itintheory in "Apple Reference Image: A New Approach for Verified Photography"]]></title><description><![CDATA[
<p>I just finished reading a scifi book called Venemous Lumpsucker which had this type of system as a minor plot point.  An interesting twist was that there were essentially smart-contract based non-disclosure agreements that could effectively disable attestation for photos and videos on a specific device that had consented to the NDA.</p>
]]></description><pubDate>Wed, 16 Sep 2026 16:00:49 +0000</pubDate><link>https://news.ycombinator.com/item?id=49729057</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49729057</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49729057</guid></item><item><title><![CDATA[New comment by itintheory in "An update on Wayback Machine access"]]></title><description><![CDATA[
<p>As someone who operates a large non-profit public data driven website, I have some VERY strong feelings about scrapers.  We looked into various commercial solutions (Datadome, HUMAN) and based on our traffic estimates from logs we'd be looking at at least 250k/yr for bot mitigation.  Anubis is offering a temporary reprieve, but after reading the recent kernel.org article [0] it's increasingly clear that this is a temporary bandaid.<p>The cheapest solution is to require a login and rate limit by API key.  I also have strong feelings about the tragedy of the commons.<p>[0] <a href="https://people.kernel.org/monsieuricon/creepy-crawlies" rel="nofollow">https://people.kernel.org/monsieuricon/creepy-crawlies</a></p>
]]></description><pubDate>Tue, 15 Sep 2026 19:09:26 +0000</pubDate><link>https://news.ycombinator.com/item?id=49717303</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49717303</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49717303</guid></item><item><title><![CDATA[New comment by itintheory in "I spent $220 on Google app ads and 60% of the installs were robots"]]></title><description><![CDATA[
<p>Looked into it - the bot operators also own, or are associated with, the sites that show the ads.  The site showing the ads gets some percentage from Google for each click through conversion, eg app install.  So app owner pays $10, google keeps $4, bot / website operator keep $6 or whatever percentages.</p>
]]></description><pubDate>Mon, 14 Sep 2026 18:15:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=49701409</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49701409</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49701409</guid></item><item><title><![CDATA[New comment by itintheory in "I spent $220 on Google app ads and 60% of the installs were robots"]]></title><description><![CDATA[
<p>How does the bot farm profit?  Did I miss it in the article somehow?</p>
]]></description><pubDate>Sat, 12 Sep 2026 11:21:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=49671204</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49671204</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49671204</guid></item><item><title><![CDATA[New comment by itintheory in "I spent $220 on Google app ads and 60% of the installs were robots"]]></title><description><![CDATA[
<p>This is the bread and butter for commercial VPN services.  The VPN knows who you are (maybe, unless you pay crypto/cash), and the rights-holders know your VPN address (from joining/logging a BitTorrent swarm, for instance), but the better providers keep no logs, and I'm guessing there's too many individuals to target effectively.<p>I've received a nasty letter from the ISP over BitTorrent (someone's phone joined my wifi, torrenting without my knowledge), but using a VPN seems sufficient.  I think the Hetzener suggestion is a VPS for running the software stack.  You could alternatively self host on hardware in your home.</p>
]]></description><pubDate>Sat, 12 Sep 2026 11:16:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=49671170</link><dc:creator>itintheory</dc:creator><comments>https://news.ycombinator.com/item?id=49671170</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49671170</guid></item></channel></rss>