Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

> If we know something of your interests (based on G+ and/or search), isn't suggesting more relevant content to you on Youtube (as a purely hypothetical example) a good thing?

No, it's not. Not because of "privacy" concerns or this whole silly "evil" debate, but because YOU'LL GUESS WRONG.

I have many interests.

I never email anyone about my search queries because I do searches about programming, and I exchange mail with my friends about past or future parties. If you try to use one to help the other you will produce a soup of irrelevant garbage.

Please just show me the most relevant links in the whole Web. I don't care what my friends think (I know already).



+googol. I just have to repeat this. Google, YOU WILL GUESS WRONG!

And my hunch is, you'll guess wrong about, oh, 8 times out of ten. Google search is the last place I go to look for something that's already one of my interests. For that I have twitter, my feed reader etc. I only google things that I have no, or little, familiarity with. Google search is supposed to tell me WHAT I DON'T ALREADY KNOW. In fact, now when I think about it, you could use my information to make search better by LOWERING rather than increasing the rank of any result that I'm likely to already know.


Use the toggle to turn off the social stuff then.


I still think it's very poor design that the toggle is:

1 - Completely unlabeled, nor is the icon indicative/clear.

2 - Not sticky. It will turn itself back on, even though this is neither implied nor stated in the UI.

3 - The only way to actually turn it off is to dig into your search settings - what percentage of your users actually know that exists, much less think about looking there?


I'd argue that a user looking to configure the settings of their search, are likely to think about looking in the search settings.


On stickiness: search settings - turn off personalized results.


A UX lesson I learned long ago: If you have to have an option for it then it's a poor design


A UX lesson I just know: It is better to have options than unchangeable defaults.

Another UX lesson: do one thing and do it well. Apparently Google forgot that. I still think it is a great one. And yes, I know it pertains to the function software plays and web empires are a different beast, but I find the poetic irony of the down the rabbit hole fiasco being related to a *nix quib delicious.


I think as a professional it's incredibly useful to have little rules like this that provide a structure out of which to make quick judgements and decisions.

But don't become complacent and let these transform into actual beliefs. There is no quicker way to kill your creativity.


Google's an AI company. If they are guessing wrong, then they're doing poorly as a company and will suffer financially. They have every incentive to make sure it guesses wrong as infrequently as possible. You have many interests, but a finite number, they can all play a role in the model of you. Similarly with people you know.

Now, if you're suggesting that for some reason there's an inherent "you will never be able to predict me with >x% accuracy", that's another topic.


> You have many interests, but a finite number

I'm not even one person. We have many computers in our home; I'm logged in to all of them, Google-wise, but they are used by my wife, the woman who looks after our kids, our children themselves, etc.

It would be very impractical to have each person log in as herself before they can search anything on Google (and my kids are too young to even be allowed to have a Google account!)

For now, customization seems to be attached only to the computer (browser), not the account; I have a computer that I use for work that no one else has access to, where Google searches work fine; I practically can't use the other computers to search Google.

If some day in the future, results are customized according to the account instead of simply the browser, then I won't be able to use Google at all.

> Google's an AI company

Ok, fine. Here's what they could do, then: detect topics in search results (and in query terms) and let one cluster or filter results by topics / domains.

Why don't they do that? Why don't they even try?


"why don't they detect and cluster by topics?"

They do - for selected queries where it makes sense.

Try searching for "jaguar" (no quotes). On the right, I have "Searches related to jaguar", which include "jaguar big cat" "jaguar car" and "jaguar download"


This. If they were actually guessing wrong, they would just stop guessing. Even if they didn't employ mostly geniuses, they would know to compare their models against random chance.

But my guess is, their models will actually do considerably better than random chance. I find google searches massively more helpful than duckduckgo searches not because ddg isn't a good search engine, but because ddg isn's snooping on me and just doesn't already know that I work a lot with such-and-such CMS.

And this is really all much ado about nothing, since if google were to ever overfit and render me unable to find something that it thinks is not my usual interest, I would just browse right back over to ddg.


I agree with you, and I think that Christopher Poole (moot) explained it very well at the Web 2.0 Summit 2011[1].

> Google and Facebook would have you believe that you're a mirror, but in fact, we're more like diamonds.

> Identity is prismatic

Those quotes may not be the best ones, but he explains well the dangers of the concept of identity according to Google and Facebook. An article about it was also featured here[2].

> If you try to use one to help the other you will produce a soup of irrelevant garbage.

I definitely agree. My different interests already enter in conflict on Youtube; and Google got trouble dealing with the fact that 90% of my searches will be in English and the last 10% will be in French, and I do not want it to filter my results towards one language or another.

1: http://www.youtube.com/watch?feature=player_embedded&v=e...

2: http://news.ycombinator.com/item?id=3123086


Think back to CW circa 1997. The internet is too complex to find the most relevant results by machine. These things must be hand-indexed. There is no better way to do it.

All Google does is guess as to what the "most relevant" results are. What you and your friends say/do/click/whatever is just another data point in the model. Presumably, it trains itself based on the results you/others actually click.

Google will almost certainly model how much a particular friend influences a particular search. It almost certainly will help, and certainly can't hurt (in the long run, if the weight goes to zero).


This.

Considering the vastly different uses people put different portions of Google's services towards, combining all of them seems rather counter-productive to me.

I suppose an option to partition different services into different relevancy 'portfolios' would solve the problem, but that puts a lot of burden on the user that, at least in my opinion, most users are not looking for.


time to switch to Bing


switch to Duckduckgo.They don't track you and it's awesome for programming questions. The first result is from Stack Overflow. You have categories on the right side. It's awesome. You can customize everything. The best search engine I ve seen in a while.


bing for main searches - DDG for tech searches... thats the play.


A query string, by itself, is just a few words in a specific order. It doesn't say very much at all. A machine has very little idea of what you're asking for. The more context it has for what you're talking about, the better chance it has at guessing what you meant -- that's all it can do, it'll never actually know.

Google's often guessing wrong now -- how many times is the link you want not the top one? How often do you have to change your query? It's all a guessing game, because the situations where there's an obvious right answer (like exact phrases) doesn't give satisfactory results. Moreover, interpreting queries for those situations can often (even usually) give sub-optimal results, as the computer's taking the query too literally.

As for programming queries, my email's actually pretty useful. I'm subscribed to a few mailing lists, and it'd be nice to not get Perl or Win32 results when I'm mostly in C and Linux. A trivial scan of my email headers would say as much.

For other interests, my youtube play list indicates a lot of what I mean. I like cats, and I watch them on youtube. When I type in jaguar, I'm actually probably asking about the cat, not the car. Lots of other people are asking about the car. It's particularly obvious if they were watching a review of a car recently, and not watching cat videos.

And as for your friends, how about the people you follow? If I'm following really good programmers, and I type in a programming-related query, I'd rather get results biased toward what really good programmers like, versus the unbiased average. Even for my friends, I actually respect some of their opinions (gasp!). If they've +1'd something in the realm of what I'm looking for, it's probably worth considering.


> The more context it has for what you're talking about, the better chance it has at guessing what you meant -- that's all it can do, it'll never actually know.

Ask me!

> I like cats, and I watch them on youtube. When I type in jaguar, I'm actually probably asking about the cat, not the car.

"Probably" but not "certainly". So what are you to do when you actually want to search about jaguar (the car)? Then you'll have to search for "jaguar car".

But that is broken.

The paradigm of (Google) search is that it gives back, as search results, documents that contain all of the search terms (implicit AND). In the query "jaguar car" the word "car" qualifies the word "jaguar" but is not necessarily to be found in result documents.

The problem with this whole guessing game is that

1) it's UNPREDICTABLE and takes control away from the user

2) it forces the user to try to "unguess" Google ("what did it try to do, and what do I have to type to make sure it guesses right") -- an exhausting mental strain, which has a tendency to fail

> If I'm following really good programmers, and I type in a programming-related query, I'd rather get results biased toward what really good programmers like, versus the unbiased average.

How can you be sure you're already following _all_ the really good programmers that have something to say about your search? Why should I have to bother with setting up all of this "following" business?

I expect Google to give me the most relevant results of the best quality, according to some general consensus -- or, if you want to go the AI route, according to actual textual analysis (for example, and at the minimum, spelling; let me filter out results where spelling mistakes are above a certain threshold, that would have value (out goes Yahoo Answers)).

Giving more weight to results coming from people I know, or even people I've already heard of, does not help.


> As for programming queries, my email's actually pretty useful. I'm subscribed to a few mailing lists, and it'd be nice to not get Perl or Win32 results when I'm mostly in C and Linux.

My email isn't very useful because I don't subscribe to mailing lists. But somehow I end up being there and I keep trying to get rid of them with the "unsubscribe" and if that doesn't work I start reporting them as spam. But the point is, essentially, no matter if you do a scan of recent emails of emails from a longer period of time there's a lot of stuff that I WAS NEVER INTERESTED IN to begin with.

More frequently I've been deleting these emails and the only emails that stay are photos (when your relatives don't quite use FB you do send photos over using Gmail) and other personal conversations. My friends and family don't program but I do, and enjoy doing so. But that never shows up in Gmail.

Even if I WAS subscribed to mailings lists, when I'm mostly Python and PHP - I would love to hear about Node.js when that suits my situation. I've nothing against learning new things. In fact, that's my second objection to you learning how I behave _now_. I WOULD LIKE TO DISCOVER NEW THINGS - DON'T TRY TO KEEP ME IN MY SHELL. (Sorry, had to use caps. I think there was a TED talk as well on this. Keep the world open - not limited - don't build walls around my current tastes to limit my future tastes.)

> For other interests, my youtube play list indicates a lot of what I mean. I like cats, and I watch them on youtube. When I type in jaguar, I'm actually probably asking about the cat, not the car. Lots of other people are asking about the car. It's particularly obvious if they were watching a review of a car recently, and not watching cat videos.

I like watching videos of lions and whatever little footage of tigers in the wild that's there. I love those kind of cats, yes. But I never really was interested in a jaguar. I like watching Top Gear but that doesn't happen on YouTube. But if I search for Jaguar, then most probably I'm looking for the car - but since you don't know what TV I watch you are going to get it wrong. Again, don't take this as a specific case but what I want to highlight is that you will never have enough information about what I consume because Google isn't the source of even, say, 70% of my consumption. I consume more from HN than from Google-originated sources for example.


"just show me the most relevant links"

Here's the rub - for a given query there is no such thing as the list of link that are "most relevant". Different people will have (potentially widely) divergent opinions on what the most relevant set of links are for the exact same query.

The goal Google web-search [1] is to show you the links from the whole web that are most relevant to you. Knowing more about you allows a search engine to make better decisions about what links are relevant to you.

You would be hard pressed to find an engineer at Google that thinks search is remotely close to a "solved" problem.

([1] for the cynics - yes, Google makes a lot of money from showing ads on search result pages. That is a nice effect of having a great search engine. Making money isn't the goal of web-search. Showing relevant results is. That was true when Larry and Sergey were cobbling together machines in their dorm room. It is still true today)


Here is a concrete example. A few days ago I visited Google and, rather tentatively, typed in the word "ring".

I was expecting to get garbage pages about wedding rings, or something like that. As its #1 result, Google sent me where I had wanted to go:

http://en.wikipedia.org/wiki/Ring_(mathematics)

By knowing my interests lie in computer science, technology, and mathematics, Google was able to return this for me as the first result.

Other searches my friends did took them to rings for sale on Overstock.com or Amazon.

I welcome this new technology.


But your own example implies that you could've achieved the same results by simply using the "and" operator: ring AND mathematics. No reason to give up privacy.

I'm actually surprised that your friends didn't get the Wikipedia entry on "ring" among their top results. When I ran the search on both Google and Bing, it came up second.


I hate to tell you this, but you'd get that as the first result even if you weren't logged in.


> The goal Google web-search [1] is to show you the links from the whole web that are most relevant to you.

But my point is: there is no such thing as "me". Now I do this (programming) and an hour later I do that (cooking). Then my daughter comes along and she wants to see pictures of puppies -- but I personally could not care less about puppies.

If Google thinks I'm into puppies and tries to infer something from that, then they will have broken search for me.

> Different people will have (potentially widely) divergent opinions on what the most relevant set of links are for the exact same query.

I really (really) don't think that's true.

What is true is that the "exact same query" could be about different things (domains), such as java the programming language, the island, or coffee.

But given a query and a domain, the most relevant links are obtained by consensus (PageRank). An interesting evolution would be to have PageRanks evaluated per topic.


But my point is: there is no such thing as "me". Now I do this (programming) and an hour later I do that (cooking). Then my daughter comes along and she wants to see pictures of puppies -- but I personally could not care less about puppies.

False problem. For one, your daughter can (and eventually will) either get her own Google account, or she will get her own computer.

But most importantly: ok, so you're not just "you", but your whole family uses the computer.

You still, as a person, get better results by Google using its knowledge of your family search habits --which include yours--, than just blindly guessing.


It is fine to show Ads depending on _that_ search. It had been a great invention of Google that had served them and us well.

But I don't want it to take my entire online life into account when I am doing a search. (At the very least I should be able to opt out of this).


I agree 100% that various levels of opting out should always be possible.

Some of that is available now - e.g. incognito mode in Chrome (or the equivalent) or simply not being logged in, Search+ has the toggle for whether or not to include personalized results are the examples that come to my mind first. One could imagine other controls that might be useful or desirable.

"But I don't want it to take my entire online life into account when I am doing a search"

Personally I think that over time, the difference in the quality of the results provided by search engines that don't know anything about it's users will so much behind those from search engines that do that we will look back and wonder how we ever found anything.

But time will tell.


No offense, but you seem to be raging a bit there.

Naturally, you have many interests. We all do. Google is trying to get the best picture of what those interests are, arguably so that they can give better than chance (or whatever they had before this). As usual, I would suspect that having more information trumps having less, so I argue that SPYW is improvement.

Sorry if this post seems impolite...


You're right, this upsets me a little. It upsets me because it takes control away from me and makes me work harder.

Instead of just searching (typing words that I expect to find in result documents), I'll have to try to second-guess Google and craft my query in such a way that it guesses right.


Well said.

I hated it it when I saw search results polluted like that a few weeks ago. Now I keep myself logged out of GMail pretty much all the time.


Yeah but social relevance is soon becoming a big part of web relevance and that stuff will be sorted with time when a balance is achieved between social and other signals, the concept is sound it will just have to be subject to iteration to get it right. Until then you can use the toggle to ignore the social stuff (the globe button).




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: