Next, the researchers tried to identify syntactic rules governing the order of these prosodic patterns, which can potentially allow future language learning models to understand and use prosody. “We noticed that there are patterns that tend to appear next to each other, in pairs, in spontaneous speech,” Weinreb explains. “It’s a simple statistical system, in which the correct choice of the next unit in a sequence depends solely on the previous one. This system works well for spontaneous conversation because it requires planning only a few seconds ahead, which is just as long as short-term memory lasts.” These pattern pairs, the researchers discovered, act as simple sentences, expressing “one new idea,” so that each pair relates to a specific topic, adding a single piece of information about it – for example, referring to a fact mentioned in the conversation and providing positive feedback.
“Our study lays the foundation for the development of an automated system that will compile a ‘dictionary’ of prosody and identify its syntactic rules for every human language and for different speaker populations,” Moses says.
“Prosody can vary depending on social status, historic events and the age of the speakers, and these variations can even manifest themselves in literary works that carefully reflect spontaneous speech,” Matalon adds. “We analyzed audiobooks as part of the study and discovered that prosodic patterns are longer in scripted speech and that the simple paired syntax of spontaneous conversation has disappeared. There are other differences, too. It’s safe to assume that the aging process and the acquisition of language in childhood are also accompanied by quantifiable prosodic changes. Moreover, there is evidence that prosody is important in internal speech – the language of thought – and that we can deepen our understanding of the existing prosody of robotic voices that are produced by speech-generating devices. The model we created promises to close the gaps that emerged over the centuries in research into expression beyond words.”
A major future application of an automated prosodic dictionary might be the development of AI capable of understanding and conveying messages through the melody of speech rather than words alone. “Imagine if Siri could understand from the melody of your voice how you feel about a certain subject, what’s important to you or whether you think you know better than her,” Weinreb adds, “and that she could adapt her response to make it sound enthusiastic or sad. We already have brain implants that convert neural activity into speech for people who can’t speak. If we can teach prosody to a computer model, we’ll be adding a significant layer of human expression that robotic systems currently lack.”