<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: ccgreg</title><link>https://news.ycombinator.com/user?id=ccgreg</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Tue, 18 Aug 2026 00:55:57 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=ccgreg" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by ccgreg in "What happens when an LLM never sees material beyond fifth grade?"]]></title><description><![CDATA[
<p>Common Crawl has very little content from Reddit and 4chan -- both blocked us years ago.</p>
]]></description><pubDate>Sun, 16 Aug 2026 15:56:44 +0000</pubDate><link>https://news.ycombinator.com/item?id=49321187</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49321187</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49321187</guid></item><item><title><![CDATA[New comment by ccgreg in "GLM-5.3: Frontier coding with emergent cyber capabilities"]]></title><description><![CDATA[
<p>Common Crawl is about 1 petabyte per year, compressed. Uncompressed is 4X.</p>
]]></description><pubDate>Fri, 14 Aug 2026 22:09:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49305189</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49305189</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49305189</guid></item><item><title><![CDATA[New comment by ccgreg in "I built a 500k-domain search engine for makers in a weekend for $10"]]></title><description><![CDATA[
<p>Both IA and CCF are private foundations and accept donations.</p>
]]></description><pubDate>Fri, 14 Aug 2026 22:07:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=49305174</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49305174</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49305174</guid></item><item><title><![CDATA[New comment by ccgreg in "A year of fighting scrapers on my 1.5 million-page website"]]></title><description><![CDATA[
<p>> It's a shame that CCBot is caught in the cross fire, but that's life.<p>We're used to it. Sadly.</p>
]]></description><pubDate>Sat, 08 Aug 2026 08:05:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=49219800</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49219800</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49219800</guid></item><item><title><![CDATA[New comment by ccgreg in "A year of fighting scrapers on my 1.5 million-page website"]]></title><description><![CDATA[
<p>The author blocked CCBot even though CCBot isn't part of the high traffic problem -- apparently he trusted Cloudflare labeling us as an "AI Bot".</p>
]]></description><pubDate>Sat, 08 Aug 2026 08:04:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=49219795</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49219795</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49219795</guid></item><item><title><![CDATA[New comment by ccgreg in "TIME Is Serving AI Bots a Different Website, with Ads Built In"]]></title><description><![CDATA[
<p>That’s already a big business!</p>
]]></description><pubDate>Wed, 05 Aug 2026 18:34:12 +0000</pubDate><link>https://news.ycombinator.com/item?id=49187005</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49187005</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49187005</guid></item><item><title><![CDATA[New comment by ccgreg in "TIME Is Serving AI Bots a Different Website, with Ads Built In"]]></title><description><![CDATA[
<p>Thanks for the heads up -- this isn't popular yet, and it requires some work to avoid polluting things like the Internet Archive Wayback Machine.</p>
]]></description><pubDate>Wed, 05 Aug 2026 15:44:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=49184503</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49184503</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49184503</guid></item><item><title><![CDATA[New comment by ccgreg in "PDF Trends 2026 Q2: Analysis of 20.6M PDFs from Common Crawl"]]></title><description><![CDATA[
<p>The post says:<p>> Because Common Crawl stores only the first 1 MB of each PDF<p>That limit became 5 MB in March 2025.</p>
]]></description><pubDate>Tue, 04 Aug 2026 18:13:34 +0000</pubDate><link>https://news.ycombinator.com/item?id=49172626</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49172626</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49172626</guid></item><item><title><![CDATA[New comment by ccgreg in "Reddit Stock Collapses 23% as AI Eats Away at User Growth"]]></title><description><![CDATA[
<p>Appreciate you double-checking.</p>
]]></description><pubDate>Sat, 01 Aug 2026 22:40:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=49139230</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49139230</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49139230</guid></item><item><title><![CDATA[New comment by ccgreg in "Reddit Stock Collapses 23% as AI Eats Away at User Growth"]]></title><description><![CDATA[
<p>That isn't true. You're welcome to peruse our index to prove or disprove your claim.</p>
]]></description><pubDate>Sat, 01 Aug 2026 21:01:35 +0000</pubDate><link>https://news.ycombinator.com/item?id=49138439</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=49138439</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49138439</guid></item><item><title><![CDATA[New comment by ccgreg in "The Second Life of Sanskrit"]]></title><description><![CDATA[
<p>It's kind of interesting that you contradict much of what the article concludes, even though the article gives a lot of examples. Maybe your prediction will be true.</p>
]]></description><pubDate>Wed, 15 Jul 2026 04:07:19 +0000</pubDate><link>https://news.ycombinator.com/item?id=48916196</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48916196</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48916196</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>We do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a page we've crawled before.</p>
]]></description><pubDate>Sat, 11 Jul 2026 07:20:01 +0000</pubDate><link>https://news.ycombinator.com/item?id=48869598</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48869598</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48869598</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>Appreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD theses helped by our public web dataset.</p>
]]></description><pubDate>Sat, 11 Jul 2026 07:17:03 +0000</pubDate><link>https://news.ycombinator.com/item?id=48869583</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48869583</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48869583</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>If you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itself is inexpensive to us and the hosting is from the AWS Open Dataset Sponsorship Program. And there's no charge for downloading it.</p>
]]></description><pubDate>Sat, 11 Jul 2026 05:04:25 +0000</pubDate><link>https://news.ycombinator.com/item?id=48868920</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48868920</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48868920</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>We aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.</p>
]]></description><pubDate>Sat, 11 Jul 2026 05:00:13 +0000</pubDate><link>https://news.ycombinator.com/item?id=48868904</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48868904</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48868904</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>Common Crawl's archive has metadata that says when each record (html file) was crawled.</p>
]]></description><pubDate>Sat, 11 Jul 2026 02:57:14 +0000</pubDate><link>https://news.ycombinator.com/item?id=48868249</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48868249</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48868249</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>Common Crawl's dataset was downloaded in full 100 times in 2025.<p>We agree that it would be great if it was even more widely used.</p>
]]></description><pubDate>Sat, 11 Jul 2026 02:56:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48868244</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48868244</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48868244</guid></item><item><title><![CDATA[New comment by ccgreg in "An update on residential proxies and the scraper situation"]]></title><description><![CDATA[
<p>A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CCBot.</p>
]]></description><pubDate>Fri, 10 Jul 2026 22:15:17 +0000</pubDate><link>https://news.ycombinator.com/item?id=48866005</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48866005</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48866005</guid></item><item><title><![CDATA[New comment by ccgreg in "Free full BGP feed. IPv4 and IPv6 (2020)"]]></title><description><![CDATA[
<p>Good timing, I'm about to release that dataset.</p>
]]></description><pubDate>Sun, 31 May 2026 05:30:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48343285</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48343285</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48343285</guid></item><item><title><![CDATA[New comment by ccgreg in "Big tech's anti-labor playbook has come for Wikipedia"]]></title><description><![CDATA[
<p>Common Crawl is working hard to improve diversity in our crawl.</p>
]]></description><pubDate>Fri, 29 May 2026 01:08:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48317725</link><dc:creator>ccgreg</dc:creator><comments>https://news.ycombinator.com/item?id=48317725</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48317725</guid></item></channel></rss>