Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

robots.txt


robots.txt is a good example of something that was introduced to make the new search engine technology more ethical but I'm not sure what point you are trying to make. If you're saying that we need a new robots.txt option saying "don't use clickstream data to help people find this page", I don't disagree; I'm not sure how many sites would take advantage of it though.


The point is that the creators of the data were given a say in how it was used. Googlers have busted a gut to provide 99% of the value of the clickstream data so they should determine how that data is used.

There's a more general point here: If something is difficult to do but easy to copy, then society should prevent that copying so that the creator can be rewarded for their effort. This maximises creativity and productivity and boosts GDP. Musicians, inventors, artists, authors, drug companies and software developers all rely on this principle. If we didn't have copyright and patents and robots.txt then there would be no incentive to produce half of the things people want, and we would all be much worse off. Google (and any other search engine) should be subject to the same rules.


You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the standards the web has already agreed upon.


To others it's just as plain that Microsoft's behavior is ethical. I'm still waiting to hear more before making up my mind. I don't blame you for being upset and frustrated, and can see why you think it's unethical. But a lot of people see it differently.

It would be great for Google and Microsoft to both come clean about all the different factors they use in their search but I'm not holding my breath.


Absolutely a new standard is needed. This is the first time any major player has suggested that data from a users clickstream should be covered by robots.txt or anything like it.

The comparison to a web directory is perfect - in the absence of a robots.txt Google and other crawlers think the data is absolutely fair game. There isn't yet a corresponding analogy for clickstream data so its fair game. Are you sure all of Google's tools respect robots.txt for passive analysis of user initiated actions on other websites?


I was using "standard" in two different ways there. Let me fix that:

>You're addressing the letter of the "law" not the spirit. Do we really need a new formal standard to indicate that this is unethical? To me it's pretty plain from the ethical standards the web has already agreed upon.


It's not plain at all. Perhaps this is a good way for the debate about clickstreams and robots.txt to start, but its certainly not settled.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: