As an Amazon Associate we earn from qualifying purchases.
So I saw an ELI5 about how "voice Recognition Works in SmartHomeHarmony Devices". It got me thinking about how ubiquitous voice control has become but how little I actually understand about the process. From what I gathered, these devices basically listen constantly for a "wake word" like "Harmony" or something similar. Once heard, they record a snippet of audio, send it to a server (probably Amazon Web services or Google cloud), which transcribes the audio to text. Then, that text is interpreted as a command like "turn on the living room lights" and sent back to the device to execute.
but what happens if the internet is down? Do all of these devices just become paperweights? and how accurate is it really? My experience is that ambient noise seems to throw it off pretty easily. I'm also curious about the privacy implications. Is that audio snippet stored somewhere permanently? It makes you wonder what data these companies are REALLY collecting and what they're doing with it all. Anyone have more insight into the limitations or vulnerabilities of voice recognition in smart home scenarios?