The Secret Origins of Amazon's Alexa
Comments
Jeff Bezos first sketched out the device that would become the Amazon Echo on a conference room whiteboard in early 2011. He wanted it to cost $20 and be controlled entirely by voice. Its brains would live in the cloud, exploiting the companyâs Web Services offerings and allowing Amazon to constantly improve it without requiring owners to upgrade their hardware.
The first-ever depiction of a device with Alexa-the artificially intelligent virtual assistant that Bezos would name after the ancient library of Alexandria-showed the speaker, a microphone, and a mute button. It wouldnât be able to understand commands right out of the box, so the sketch identified the act of configuring the device to a wireless network as a challenge requiring further thought.
Greg Hart, who was Bezosâ technical adviser, or âTA,â at the time, was the other person in the meeting, and he was listening closely. Bezos said he wanted Hart to lead the group that would turn this somewhat outlandish notion for a voice computer into an actual product. Hart snapped a photo of the drawing with his phone. âJeff, I donât have any experience in hardware, and the largest software team Iâve led is only about 40 people,â he recalls saying. âYouâll do fine,â Bezos replied. Hart thanked him for the vote of confidence and said, âOK, well, remember that when we screw up along the way.â
For the next three years, Bezos would remain intimately involved in the project. He authorized the investment of hundreds of millions of dollars before the first Echo was ever released, made detailed product decisions, and met with the team as frequently as every other day. Using the German superlative, employees referred to him as the ĂŒber product manager. But it was Hart who ran the effort, just across the street from Bezosâ office, in a building that housed the team working on the Kindle.
Building the Team and Secrecy
Over the next few months, Hart hired a small group from inside and outside the company. Like his boss, he was obsessed with secrecy. He sent out vague emails to prospective hires with the subject line âJoin my missionâ and asked interview questions like âHow would you design a Kindle for the blind?â He declined to specify what product candidates would be working on. One interviewee recalls guessing that it was Amazonâs widely rumored smartphone and says that Hart replied, âThereâs another team building a phone. But this is way more interesting.â
The initial Alexa crew worked with a feverish sense of urgency. Unrealistically, Bezos wanted to release the device in six to 12 months. He would have a good reason to hurry. On October 4, 2011, just as the Alexa team was coming together, Apple introduced the Siri virtual assistant in the iPhone 4S. It was the last passion project of cofounder Steve Jobs, who died of cancer the next day. Hart and his team felt validated by the news that a resurgent Apple was also working on a voice-activated personal assistant, but they were discouraged by the fact that Siri was first to market and initially garnered some negative reviews.
The Amazon team tried to reassure themselves that their product was unique, since it would be independent from smartphones. They were also attempting to pull off a much more technically complex feat. Siriâs users spoke commands directly into microphones. Amazon was trying to build a service capable of understanding language spoken from across a noisy room, using a relatively immature technology called far-field speech recognition.
Acquiring Speech Technology
To speed up development, Hart and his crew went looking for startups to acquire. It was a nontrivial challenge, since Nuance, the Boston-based speech giant whose technology Apple had licensed for Siri (and which was recently acquired by Microsoft), had grown over the years by gobbling up the top American speech companies. Alexa execs tried to learn which of the remaining startups were promising by asking prospective targets to voice-enable the Kindle digital book catalog, then studying their methods and results. The search led to several rapid-fire acquisitions over the next two years, including the Polish startup Ivona.
Ivona was founded in 2001 by Lukasz Osowski, a computer science student at the GdaĆsk University of Technology. Osowski had the notion that so-called text-to-speech, or TTS, could read digital texts aloud in a natural voice and help the visually impaired in Poland. With a younger classmate, Michal Kaszczuk, he took recordings of an actorâs voice and selected fragments of words, called diphones, and then blended or âconcatenatedâ them together in different combinations to approximate natural-sounding words and sentences that the actor might never have uttered.
The Ivona founders got an early glimpse of how powerful their technology could be when they paid a popular Polish actor named Jacek Labijak to record hours of speech to create a database of sounds. The resulting product, which they called Spiker, quickly became the top-selling computer voice in Poland. Over the next few years, it was used widely in subways, elevators, and for robocall campaigns. Labijak subsequently began to hear himself everywhere, and regularly received phone calls in his own voice urging him, for example, to vote for a candidate in an upcoming election. Pranksters manipulated the software to have him say inappropriate things and posted the clips online, where his children discovered them. The Ivona founders then had to renegotiate the actorâs contract after he angrily tried to withdraw his voice from the software. (Today âJacekâ remains one of the Polish voices offered by AWSâ Amazon Polly computer voice service.)
In 2006, Ivona began to enter and repeatedly win the annual Blizzard Challenge, a competition for the most natural computer voice, organized by Carnegie Mellon University. By 2012, Ivona had expanded into 20 other languages and offered more than 40 voices.
Hart and Al Lindsay, the first engineering manager on the project, visited them in GdaĆsk on a trip they were taking through Europe to look for acquisition targets. âFrom the minute we walked into their offices, we knew it was a culture fit,â Lindsay says, pointing to Ivonaâs progress in a field where researchers often get distracted by high-minded pursuits and have a difficult time shipping actual products. âTheir scrappiness allowed them to look outside pure academia and not be blinded by science.â The purchase, for around $30 million, was completed in 2012 but kept secret for a year. The Ivona team and the growing number of speech engineers Amazon would hire for its new GdaĆsk R&D center were put in charge of crafting Alexaâs voice.
Crafting Alexaâs Voice
The program was micromanaged by Bezos himself and subject to the CEOâs usual curiosities and whims. At first, Bezos said he wanted dozens of distinct voices to emanate from the device, each associated with a different goal or task, such as listening to music or booking a flight. When that proved impractical, the team considered lists of characteristics they wanted in a single personality, such as trustworthiness, empathy, and warmth, and determined those traits were more commonly associated with a female voice.
To develop this voice and ensure it had no trace of a regional accent, the team in Poland worked with an Atlanta-area-based voice-over studio, GM Voices, the same outfit that had helped turn recordings from a voice actress named Susan Bennett into Appleâs agent, Siri. To create synthetic personalities for its customers, GM Voices gives voice actors hundreds of hours of text to read, from entire books to random articles, a mind-numbing process that could stretch on for months.
Believing that the selection of the right voice for Alexa was critical, Hart and colleagues spent months reviewing the recordings of various candidates that GM Voices produced for the project, and they presented the top picks to Bezos. The Amazon team ranked the best ones, asked for additional samples, and finally made a choice. Bezos signed off on it.
Characteristically secretive, Amazon has never revealed the name of the voice artist behind Alexa. I learned her identity after canvasing the professional voice-over community: voice actress and singer Nina Rolle, who is based in Boulder, Colorado. Her professional website contains links to old radio ads for products such as Mottâs Apple Juice and the Volkswagen Passat-and the warm timbre of Alexaâs voice is unmistakable. Rolle said she wasnât allowed to talk to me when I reached her on the phone in February 2021. When I asked Amazon to speak with her, they declined.
Beta Testing and Bezosâs Frustration
Alexa now had a voice, but it soon became clear that she needed a new brain. In early 2013, Amazon began moving a prototype of the original Echo into the homes of hundreds of employees, who were asked to sign confidentiality agreements and fill out surveys about their experiences with the product. The experimental devices were, by all accounts, slow and dumb. Perhaps the most harrowing review came from Bezos himself. The CEO was apparently testing a unit in his Seattle home, and in a pique of frustration over its lack of comprehension, he told Alexa to go âshoot yourself in the head.â One of the engineers who heard the comment while reviewing interactions with the test device said, âWe all thought it might be the end of the project, or at least the end of a few of us at Amazon.â
In the months that followed, Amazonâs ongoing efforts to make its product smarter would become embroiled in a battle between dueling AI dogmas and would lead to its biggest challenge yet.
The Battle of AI Methods
Thanks to the acquisition of an artificial intelligence company in Cambridge, England, called Evi, Alexa was already proficient in the culturally common chitchat called phatic speech. If a user said to the device, âAlexa, good morning, how are you?â Alexa could make the right connection and respond. It could also handle factual queries, such as requests to name the planets in the solar system. These qualities, the result of a programming technique called knowledge graphs, gave the impression that Alexa was smart. But was it?
Proponents of another method of natural language understanding, called deep learning, believed that Eviâs method was too regimented to give Alexa the kind of authentic intelligence that would satisfy Bezosâ dream of a versatile assistant that could talk to users and answer any question. If a user said, âPlay music by Sting,â for instance, they feared a knowledge-graph-based system could think he was trying to say âbyeâ to the artist and get confused. In the deep learning method, machines were fed large amounts of data about how people converse and what responses proved satisfying, and then were programmed to train themselves to offer the best answers. In other words, the more Alexa was used, the smarter it would get.
The chief proponent of this approach was an Indian-born engineer named Rohit Prasad. Prasad and his colleagues had to solve the paradox that confronts all companies developing AI: If they launch a system that is dumb, customers wonât use it, and therefore wonât generate enough data to improve the service. But companies need that data to train the system to make it smarter.
Google and Apple solved the paradox in part by licensing technology from Nuance, using its results to train their own speech models and then afterward cutting ties with the company. For years, Google also collected speech data from a toll-free directory assistance line, 800-Goog-411. Amazon had no such services it could mine, and Hart was against licensing outside technology-he thought it would limit the companyâs flexibility in the long run. But the meager training data from beta tests in employeesâ homes amounted to speech from a few hundred white-collar workers, usually uttered from across a noisy room in the mornings and evenings when they werenât at the office. The data was lousy, and there wasnât enough of it.
Meanwhile Bezos grew impatient. âHow will we even know when this product is good?â he kept asking. Hart, Prasad, and their team created graphs that projected how Alexa would improve as data collection progressed. The math suggested they would need to roughly double the scale of their data collection efforts to achieve each successive 3 percent increase in Alexaâs accuracy.
The Pivotal Meeting
That spring, only a few weeks after Prasad had joined the company, the team brought a six-page narrative to Bezos that laid out these facts, and they proposed to double the size of the speech science team and postpone a planned launch from the summer into the fall. The meeting did not go well. âYou are going about this the wrong way,â Bezos said after reading about the delay, according to someone who was present. âFirst tell me what would be a magical product, then tell me how to get there.â
Bezosâ technical adviser at the time, Dilip Kumar, then asked if the company had enough data. Prasad, who was calling into the meeting from Cambridge, replied that they would need thousands of more hours of complex, far-field voice commands. According to an executive who was in the room, Bezos apparently factored in the request to increase the number of speech scientists and did the calculation in his head in a few seconds. âLet me get this straight. You are telling me that for your big request to make this product successful, instead of it taking 40 years, it will only take us 20?â Prasad tried to dance around it. âJeff, that is not how we think about it.â âShow me where my math is wrong!â Bezos said, according to a person who was in the room. Hart jumped in. âHang on, Jeff, we hear you, we got it.â
Prasad and other Amazon executives would remember that meeting, and the other tough interactions with Bezos during the development of Alexa, differently. But according to a person who was there, the CEO stood up and said, âYou guys arenât serious about making this product,â and abruptly ended the meeting.
After Jeff Bezos walked out on them, the Alexa executives working on the prototype retreated with their wound.
Comments
No comments yet. Start the discussion.