Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

(disclaimer: this is based on demonstrations I have seen - I don't have access to JB and thus I can only speak of Siri):

In many of these demonstrations I noticed one thing that was bugging me: Even though the voice recognition in Android seems really cool (I don't have JB yet), it doesn't give definite audible confirmation of the command in many cases and sometimes it even requires user interaction with the screen.

Now personally, I already believe that speech input is kind of a gimmick in itself (try using english voice recognition with my address book filled with german names...), I believe that to even have a chance to move from gimmick to useful feature, it must work without user interaction on the screen.

"Play <whatever band name>" followed by "beep" and the a button to press on the screen doesn't help me. A useful response to "play <insert band name here>" is "playing <insert band name here>" followed by actually playing it.

Or "call <some name>" - if you just get back "calling" or even just a beep - how would you know whether the recognition was successful or not and the correct name has been recognized?

Some commands on Android seem to be doing fine (the weather example), but others fail in one way (the play example seems to require user interaction on the screen) or another (the "turn on wifi" command doesn't produce any audible confirmation or error message - just the same beep sound as if it worked).

Siri, while it might not have as good a recognition as the Android solution, is much better in that regards: It always confirms your command. As such Siri seems moderately more useful as an additional input method whereas Android, by forcing you to look at the screen when inputting a voice command, reduces this to a gimmick and nothing else.



> ... it doesn't give definite audible confirmation of the command ...

I don't have this either, but can tell you that Google's voice recognition does have a confidence rating on a per word basis. (See Google voice messages. Also use the recognition API and it provides alternatives.)

In their navigation product they directly do whatever was said if there is a high confidence. If not it shows the recognized speech with a "pie" based countdown to using the displayed recognition. You can press OK to go ahead (or wait), or cancel/try again.

They could obviously do something similar with this.

They also use context for their voice recognition. I grew up in a town named Piggs Peak (note two 'g's). If you say "piggs peak" you will get that spelling, but for example saying "peak pigs" gets you the spelling with only one g. This explains why wooster/worcestor doesn't confuse them. I don't have a siri capable device so I don't know what they do.


It appears that for queries that don't have direct "answer" type response and will perform an action it displays a progress bar with the query it understood. Presumably this is so that you can cancel the action if voice recognition was wrong.

It doesn't require any input though - I just tested it. Once the progress bar reaches the end (it seems to take ~7 seconds) it will complete the action.


That means that in case it mis-understood you that you have to wait ~7 seconds before you notice that it was wrong.

If it immediately confirmed like Siri, you would know right then and could re-issue the command.


Are you trolling? I understand you said in your parent that you don't have JB and are strictly going by the video and demo.. but how could you miss this? It's in the first minute.. multiple times.

You do not have to "wait ~7 seconds before you notice that it was wrong" so you can "re-issue the change". The result is displayed immediately. You can cancel the auto-action (which you first said didn't even exist), or force it through before the ~5 seconds (not ~7) elapses.

Go re-watch the entire video in the foreground, please.


I think pilif meant that you'd have to wait for the progress bar (about seven seconds) to finish before you noticed something was wrong. For example:

  "Call the Drake Hotel in Toronto."
  *bling* "Calling..."
  (wait seven seconds)
  "Hey, this is Drake. What's up?"
Versus what Siri does:

  "Call the Drake Hotel in Toronto."
  *bling* "Calling Drake Smith..."
  "No, wait! Stop!"
Think about using the voice commands when you can't see the device. Like when you're driving or running. It's useful to have the audible feedback in addition to whatever's displayed on the screen.


It works the same way as Voice Search has always worked (at least since Froyo): The found action is shown to you for about 3 seconds along with a timed progress meter (3 seconds?) and buttons to proceed or cancel. When the timer ends, the action proceeds.

This is the correct solution, IMO. It would be quite frustrating to have the wrong phone number instantly begin to dial, for instance. One time when I said "call <name of restaurant>", it came up with "Call <name of restaurant>" with the address of the location I didn't want shown beneath. This gave me time to tap Cancel, which then showed me a list of the alternative results/locations.


So lets say you have your phone in your bag, not looking at it. Then you enable the voice command and say "Call Foo Burgers"

Your phone understands this as "Call Bar Burgers" and shows on the screen "Calling Bar Burgers". The phone makes a "beep" sound and then proceeds to show a progress bar which you don't see because your phone is in your pocket.

Then the phone connects and you learn of your mistake as the person at the other end answers with "This is Bar burgers, Mr. Foobar speaking".

The only way around this is to enable the voice command, take the phone out of your pocket and then check what it says above the progress bar.

With siri, if you say "Call Foo Burgers", Siri would respond (in audio over your headphones) with "Calling Bar Burgers", giving you a chance to cancel before you annoy the person at the other end and without forcing you to take the phone out of your pocket to check (which is the point of voice commands)


I already have to hit a button to enter a voice a command, I dont mind at all hitting another one to accept it and complete the deal.

I use Voice commands for pretty much all input to Google Maps and Navigation (and nowhere else). That it works flawlessly in my experience even with my mumbling is plenty good enough.

The point of voice commands, when I use them, isn't to avoid any visual interaction, it's just so I don't have to type something I don't want to type.


> Siri, while it might not have as good a recognition as the Android solution, is much better in that regards: It always confirms your command.

Even better, Siri gives confirmation by default, but you can disable audible feedback if you so prefer.


If you look closely, after it understands "play" command, there's a progress bar. If the user doesn't respond it would go to the music player automatically. He cut it short and clicked "Play", probably for TL;DW people.


You are exactly correct; It does this for anything that may require an action. The action can be cancelled prior to the progress bar completing, but none-interaction will result the progress bar completing and the music playing, alarm being set, etc..

I don't understand how parent can write such a lengthy comment without watching the entire video and understanding what they're talking about. Even the parent's disclaimer states they are only going by the video, but it clearly shows they didn't watch it fully, since what they missed is contained within the very first minute of the video, multiple times.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: