Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I didn't say the ordered list was obvious. But it is an ordered list (at least to the user -- it may not be ordered all the way down or partially so).

Fair enough point about suggest and instant. In those cases Google is adding value to the query, but I think we'd both agree that it is incrementally so. Instant generally just helps you get to the eventual query faster. Spell checking does add value, but its a +epsilon. In either case, your right, we should add this to our function.

In any case I think my point still holds. The query is the most important aspect. The browsed page, for indexing purposes is the second most important aspect. The association between the query and the URL is the least important aspect.

IOW, the least value for Bing is the search algorithm. The query and page to add to the index (as a browsed page) are the most important.



>The association between the query and the URL is the least important aspect.

I don't understand why you believe this. This is the entire basis for why search is hard. It is the hard problem that search tries to solve. Many companies have spent in aggregate, billions of dollars in R&D trying to solve this, and all but a few have folded. It's an "AI-complete" problem in that solving it perfectly would be a sufficient demonstration of strong AI. It's the whole reason why Google exists.


The reason is that with the URL and the query, there's a decent chance you can derive the association.

Now the reason I say this is that to my best approximation the key difference that makes Google better than Bing (when it is) is the index size, and not the search algorithm.

So the two key pieces of value are: A) queries that people don't do on your site, which would have generated poor relevance

B) Pages that are actually browsed that you don't have indexed or up to date.

Now don't get me wrong. There is value in the association, but I think I'd capture 80% of the value with the above. Now as you point out, these lists are often non-obvious, but Bing does a near equal job of creating the lists. And if you give them the search terms where they have gaps. And fill the index, then I think the gap closes most of the way.

And lets be clear, because they clicked the link doesn't mean its a good link. See ehow.com or expertsexchange. But it does have some value.

And frankly I think MS would have been willing to let it go if Google had gone straight to MS with it and said, "We have this data. Even though its not unethical we can spin it to the media to make it look bad. Kill the association." MS probably would have as the net win isn't that huge. Now the PR loss in pulling it would be worse than the PR loss in keeping it (with no integrity loss, since they [and I] think it is perfectly ethical).


>is the index size, and not the search algorithm

I'm pretty thoroughly sure this is false. I'm sorry I can't show data to back it up, but Google has many many systems that are years ahead of bing's technology. I'm kinda hamstrung here by being unable to reveal anything about them. To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their indexes. Bing is just unable to return this data for as many queries as Google is.

>but Bing does a near equal job of creating the lists.

To the extent that this is true, how much of it is due to data harvested from Google search? This is something that google can't really demonstrate with evidence, and what I would have hoped that bing would clarify in a public statement if they had anything defensible to say.


To a decent approximation, both Google and Bing likely have the whole internet that matters (and that they are allowed to crawl) in their indexes. Bing is just unable to return this data for as many queries as Google is.

Maybe this is true, but its not apparent. I had commented on this a month or so back that I thought Bing was better at "vague" queries, where I don't know exactly what I'm looking for. But I'll know it when I see it.

Whereas Google is really good at targeted queries. I need info on the HP battery model number 003D434F90. These searches in Bing will often bring back literally nothing, while Google will often have one or two links, but they happen to be the link that I want. The text of the query is almost always found in these pages.

From that I infer it is index size, since the text is in the page.

To the extent that this is true, how much of it is due to data harvested from Google search?

While I find the quality of searches similar I don't find the results to be similar, if that makes sense. If Bing is harvesting your results, they're still using other very clever methods to surface other equally good, yet different results.


One obvious and very public example of how the algorithm matters more than the index size would be Cuil. It was launched with lots of hype about how it would have a index several times larger than Google's. And of course the results were notoriously bad.

I'm stunned that anyone could think that the core issue in this controversy is indexing or page discovery. It should have been totally obvious from the examples in the initial article and blog post that this was specifically about copying ranking.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: