Will Nutch be a distributed, P2P-based search engine?
We don’t think it is presently possible to build a peer-to-peer search engine that is competitive with existing search engines. It would just be too slow. Returning results in less than a second is important: it lets people rapidly reformulate their queries so that they can more often find what they’re looking for. In short, a fast search engine is a better search engine. I don’t think many people would want to use a search engine that takes ten or more seconds to return results. That said, if someone wishes to start a sub-project of Nutch exploring distributed searching, we’d love to host it. We don’t think these techniques are likely to solve the hard problems Nutch needs to solve, but we’d be happy to be proven wrong.