<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Hacker News: n1xis10t</title><link>https://news.ycombinator.com/user?id=n1xis10t</link><description>Hacker News RSS</description><docs>https://hnrss.org/</docs><generator>hnrss v2.1.1</generator><lastBuildDate>Thu, 27 Aug 2026 10:27:29 +0000</lastBuildDate><atom:link href="https://hnrss.org/user?id=n1xis10t" rel="self" type="application/rss+xml"></atom:link><item><title><![CDATA[New comment by n1xis10t in "Ask HN: Something fishy is going on with Google feed on my Android"]]></title><description><![CDATA[
<p>What does the icon look like?</p>
]]></description><pubDate>Mon, 27 Jul 2026 01:03:42 +0000</pubDate><link>https://news.ycombinator.com/item?id=49064011</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=49064011</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49064011</guid></item><item><title><![CDATA[New comment by n1xis10t in "Ask HN: Something fishy is going on with Google feed on my Android"]]></title><description><![CDATA[
<p>Such a weird thing. .aigc isn’t a real toplevel domain, and when I searched for it I found that apparently TikTok uses it as an abbreviation for “ai generated content”. That fits with the fake headlines. I was going to suggest checking /etc/hosts to see if some lines had been put in there by malware to make fake domain names capable of resolving, but you say the links don’t work. So that’s weird.<p>So is this just something wrong with Google?</p>
]]></description><pubDate>Mon, 27 Jul 2026 01:03:10 +0000</pubDate><link>https://news.ycombinator.com/item?id=49064007</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=49064007</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=49064007</guid></item><item><title><![CDATA[New comment by n1xis10t in "So I am building my own web search engine"]]></title><description><![CDATA[
<p>Wow, so ~43 million a day, ~15 billion a year? That’s sick.<p>How much space do you have on the server, and how many pages do you intend to scale to?</p>
]]></description><pubDate>Tue, 14 Jul 2026 05:09:29 +0000</pubDate><link>https://news.ycombinator.com/item?id=48902530</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48902530</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48902530</guid></item><item><title><![CDATA[New comment by n1xis10t in "Ask HN: Can anyone explain this Gsearch rabbit-hole?"]]></title><description><![CDATA[
<p>Yeah that’s pretty weird. I thought that maybe it was websites getting hacked and then redirecting to malicious stuff, but then I did your recommended search and it all seems very benign, and the videos seem to match the page title and urls. So I thought maybe it was some service that people use to redirect to youtube but have their own link instead of a youtu.be link, but a search for "hook global" "youtube" "redirect" doesn’t appear to reveal what this is, it just shows more examples of the weirdness. Including pages that are apparently just from youtube. Maybe if you turn off redirection in your browser and then take a look at the source for the html page that redirects you, it’ll give a clue. I can’t right now (on mobile) but I can probably try later.</p>
]]></description><pubDate>Mon, 13 Jul 2026 19:55:30 +0000</pubDate><link>https://news.ycombinator.com/item?id=48897913</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48897913</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48897913</guid></item><item><title><![CDATA[New comment by n1xis10t in "Ask HN: Does your mind drift while waiting for AI prompts to finish?"]]></title><description><![CDATA[
<p>What will happen if you don’t use AI? Will they fire you?</p>
]]></description><pubDate>Wed, 17 Jun 2026 12:55:04 +0000</pubDate><link>https://news.ycombinator.com/item?id=48569811</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48569811</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48569811</guid></item><item><title><![CDATA[New comment by n1xis10t in "Ask HN: Why is packages.ubuntu.com not being indexed by Google?"]]></title><description><![CDATA[
<p>Was the reason for your attempt to search in this way that you got the “internal server error” message (that I just got) when you tried to use the website’s search feature?</p>
]]></description><pubDate>Fri, 12 Jun 2026 16:54:16 +0000</pubDate><link>https://news.ycombinator.com/item?id=48506500</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48506500</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48506500</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: I built a social-first search(X, Reddit, TikTok) cause Google is boring"]]></title><description><![CDATA[
<p>It seems pretty cool, but I don’t think I’m the best person to ask. Most of social media I don’t use, and I’m not terribly interested if the results are from the apis directly. If it was an independent index of these sites, or if it took a really long time to get all the available results from the apis and then reranked them, then I would be interested because there is a lot of censorship and downranking, and it would be nice to get past that.<p>I just tried it with the query “reckless ben”, and in the “all” tab a lot of the results are from various websites (like kotaku), where do these results come from?</p>
]]></description><pubDate>Thu, 11 Jun 2026 14:02:31 +0000</pubDate><link>https://news.ycombinator.com/item?id=48490500</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48490500</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48490500</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: I built a social-first search(X, Reddit, TikTok) cause Google is boring"]]></title><description><![CDATA[
<p>How does it work? Is it like a meta search engine where your service sends the query to all the apis, or do you actually have an index of all this stuff?</p>
]]></description><pubDate>Tue, 09 Jun 2026 18:03:39 +0000</pubDate><link>https://news.ycombinator.com/item?id=48464980</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48464980</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48464980</guid></item><item><title><![CDATA[New comment by n1xis10t in "Google's top result is 16yo when searching for "Ubuntu 24.04 install fonts""]]></title><description><![CDATA[
<p>I haven’t tried that, thanks for the tip. It is far from perfect of course, because there are tons of great high quality sites that use .com. What we really need is a good search engine that is the scale of google, but those don’t happen along very often. Cuil was really big, but they really messed it up.</p>
]]></description><pubDate>Mon, 01 Jun 2026 14:37:46 +0000</pubDate><link>https://news.ycombinator.com/item?id=48357426</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48357426</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48357426</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: My independent search engine focused on user control"]]></title><description><![CDATA[
<p>Gotcha, it’ll be interesting to see how it progresses.</p>
]]></description><pubDate>Fri, 22 May 2026 22:51:08 +0000</pubDate><link>https://news.ycombinator.com/item?id=48242609</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48242609</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48242609</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: My independent search engine focused on user control"]]></title><description><![CDATA[
<p>Oh, the rules are that paged out articles are only one page, so a longer article would have to go somewhere else like 2600.<p>The collection of about 1.073 million pages in extracted text form that I have takes up about 4.8 GiB spread across 15 files (but compressed they’re only 2.1 GiB), so if you were just downloading them until your hard drive filled up you’d have about 107 million pages, and you’d need something like 5TiB for a billion pages. These are the WET files from CC, which are extracted text only. I know the WARC files are made so that if you know the correct offset in bytes, you can take out individual documents without decompressing the whole file, but I’m not sure if the WET ones work the same way. If they do, your pile of text could be a bit less than half the size and still usable with an index.<p>I don’t know how much more space an index of the data would take up, but I think it really depends on how complicated it is. If the index is super basic, like “give me a keyword and I’ll give you a list of docs that it appears in”, then I think the index should be smaller than the text collection. You use embeddings and stuff, so I don’t know how big it would be.<p>Marginalia search has about 1 billion pages, and when someone asked how big the index is on disk, he said this: “16 TB for the unprocessed crawl data (compressed). 7.7 TB for the files that actually constitute the index (positions data, reverse index)” I’m guessing that the “unprocessed crawl data” is raw html, and that’s why it’s significantly larger than my Blekko-era Common Crawl extracted text based estimate.<p>So with an uncompressed pile of extracted text and a Marginalia style index, one billion pages would be about 13 TB on disk. He says “positions data” though, so I think that means that the locations of keywords in documents is part of the index. Probably the original extracted text and the position data don’t both need to be there (and they’re probably about the same size), so you would just pick between having the original documents and needing to use compute to find the keyword positions for ranking, or having the keyword positions for ranking and needing to use compute to reconstruct the original documents (if needed). So if you pick one instead of having both, the whole thing probably just takes up about 7.7TB.<p>Oh also, downloading these files from the Common Crawl should go really fast. One file has about 73000 documents in it, and takes up around 141 MiB (in it’s compressed form, but that’s the form it’ll be downloaded in.)<p>These wouldn’t get you recent stuff of course, but they would make the index size way bigger, and so the quality would go up but it would be dated. It would be like resurrecting Blekko. For context, Greg Lindahl said that their largest index was 4 billion pages, but that their crawl frontier was much larger.<p>Here’s another idea: Download tons of old stuff from the Common Crawl / Blekko, but only keep and index the pages that are inaccessible today. This would make your search engine as competitive [edit: probably complementary is a better word] as possible, because it draws from resources that the other engines don’t have. I’m pretty sure the standard is to prune 404’s from search indexes, which seems very silly to me because cached page content can be served, or a link to the Wayback machine can be given. I suppose there are a couple partial exceptions, because Kagi, Brave, and either yep.com or Yandex will give some results from the wayback machine, but I imagine this is a very small part of what they have.</p>
]]></description><pubDate>Fri, 22 May 2026 17:14:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48238667</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48238667</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48238667</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: My independent search engine focused on user control"]]></title><description><![CDATA[
<p>You’re Canadian? That’s pretty hilarious, I am too. It must be something they put in the Timbits.<p>For promotion, I’d recommend picking the most technically interesting part of your implementation, something that’s really clever, and then making a one page writeup for Paged Out magazine about it (<a href="https://pagedout.institute/" rel="nofollow">https://pagedout.institute/</a>). They regularly have interesting stuff to read, and they have a pretty decent amount of readers. You could write something longer and send it in to 2600 magazine too, they’d probably be interested even if it was an overview of the project.<p>Maybe the engine should be bigger first though so people are more enthused when they try it. I think 1 billion pages is around where a search engine starts to seem more normal: that’s about how much Marginalia has. How much space on disk does your index take up right now? Would you say the bottleneck is more the hard drive space or the crawling speed?</p>
]]></description><pubDate>Fri, 22 May 2026 02:54:36 +0000</pubDate><link>https://news.ycombinator.com/item?id=48231432</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48231432</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48231432</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: My independent search engine focused on user control"]]></title><description><![CDATA[
<p>Very cool, I subscribed to the newsletter. I’ve experimented with retrieval and ranking across a sample of a million pages from the early days of the Common Crawl (around 2014) and I was surprised by how many of them seemed high quality. The CTO of CC tells me it’s because most of the early URLs were donated by Blekko, which was an old search engine that he used to work for. I don’t know what the quality of recent CC stuff is like, but I think it would be fun to supplement an index with this older data, especially because you’d get a lot of pages that are 404’s now (but you could deliver the extracted text to the user, or link to a temporally nearby snapshot from WayBack).<p>Another fun thing to consider is making a meta search engine that functions like MetaCrawler used to, where it gets all (or a bunch of) the available results from all the source engines, and then actually fetches and extracts the text from the linked pages, and then matches the query and ranks the pages independent of what the source engines did. If you’d like to do that, I would recommend adapting the source code of 4get.ca (at least for the scrapers), because the guy who writes it is rather talented at coming up with and maintaining workarounds.<p>If you monetize this, I’d be interested in working for you. I know Python, HTML, CSS, am familiar with JavaScript, and have a lot of experimental (and successful!) experience with ranking web results.<p>Also, you might be interested in reading this article (from 2600 magazine) about disappearing search engines: <a href="https://archive.org/details/search-timeline" rel="nofollow">https://archive.org/details/search-timeline</a> In addition to the things in that article, there was a search engine for discord (“Searchcord”) that went away in less than a week after it was announced here (on HN), and there is this recent blog post which lists search engines with independent indexes, a painfully large number of which went away with no announcement: <a href="https://seirdy.one/posts/2021/03/10/search-engines-with-own-indexes/" rel="nofollow">https://seirdy.one/posts/2021/03/10/search-engines-with-own-...</a> The author of the 2600 article doesn’t really get into theories about why search engines disappear, but it certainly seems like a lot of them do. I’m curious to know if they disappear for random different reasons, or if it’s just really difficult to make and maintain a search project, or if there’s some other common reason. If you suddenly feel disinclined to work on this project, could you let me know why (maybe anonymously with a new email account or something)? Thanks.</p>
]]></description><pubDate>Thu, 21 May 2026 18:24:54 +0000</pubDate><link>https://news.ycombinator.com/item?id=48226990</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48226990</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48226990</guid></item><item><title><![CDATA[New comment by n1xis10t in "TheArchiveBase – A Lightweight Distributed Search Engine"]]></title><description><![CDATA[
<p>Would you mind writing a comment about this search engine you have found or created? I’m intrigued, but I don’t like clicking on links right away.</p>
]]></description><pubDate>Tue, 19 May 2026 02:58:59 +0000</pubDate><link>https://news.ycombinator.com/item?id=48188678</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48188678</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48188678</guid></item><item><title><![CDATA[New comment by n1xis10t in "TheArchiveBase – A Lightweight Distributed Search Engine"]]></title><description><![CDATA[
<p>Why?</p>
]]></description><pubDate>Tue, 19 May 2026 02:57:48 +0000</pubDate><link>https://news.ycombinator.com/item?id=48188671</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48188671</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48188671</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: A search engine for deleted YouTube videos (1.5B+ indexed since 2005)"]]></title><description><![CDATA[
<p>I just thought to try putting “music” in place of “www” in the playlist url, but unfortunately IA and CC still have nothing.</p>
]]></description><pubDate>Wed, 13 May 2026 15:56:09 +0000</pubDate><link>https://news.ycombinator.com/item?id=48123602</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48123602</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48123602</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: A search engine for deleted YouTube videos (1.5B+ indexed since 2005)"]]></title><description><![CDATA[
<p>Unfortunately I got the same result that archivarix did above, nothing in CC and nothing in IA. The thumbnail had a different link than the title of the playlist so I thought I’d try that, but the wayback machine redirects to the first video on the playlist, and there’s nothing in the common crawl. I also checked the html just in case there was a track listing hidden in there, but there wasn’t.</p>
]]></description><pubDate>Wed, 13 May 2026 15:42:33 +0000</pubDate><link>https://news.ycombinator.com/item?id=48123408</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48123408</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48123408</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: A search engine for deleted YouTube videos (1.5B+ indexed since 2005)"]]></title><description><![CDATA[
<p>I might not be able to help, but I’d like to give it a shot. Can I see the two IA links? If it was the “this page hasn’t been archived” error I’m less likely to be able to do anything, but I can check all the Common Crawl indexes and see if they have a copy of something useful.</p>
]]></description><pubDate>Wed, 13 May 2026 01:48:27 +0000</pubDate><link>https://news.ycombinator.com/item?id=48116897</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48116897</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48116897</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: A search engine for deleted YouTube videos (1.5B+ indexed since 2005)"]]></title><description><![CDATA[
<p>Update: So I mustered the courage to try the search engine, because it was looking not very much like a scam, and it becomes very apparent as soon as you use it that non-deleted videos are also indexed.</p>
]]></description><pubDate>Sun, 10 May 2026 14:23:28 +0000</pubDate><link>https://news.ycombinator.com/item?id=48084228</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48084228</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48084228</guid></item><item><title><![CDATA[New comment by n1xis10t in "Show HN: A search engine for deleted YouTube videos (1.5B+ indexed since 2005)"]]></title><description><![CDATA[
<p>Seems pretty cool. So this is a recent project, and you haven’t been working on this since 2005 right?<p>Have you considered also indexing videos that haven’t been deleted?</p>
]]></description><pubDate>Sun, 10 May 2026 02:26:00 +0000</pubDate><link>https://news.ycombinator.com/item?id=48080415</link><dc:creator>n1xis10t</dc:creator><comments>https://news.ycombinator.com/item?id=48080415</comments><guid isPermaLink="false">https://news.ycombinator.com/item?id=48080415</guid></item></channel></rss>