
AI can generate more natural synthesized speech by including pauses
shutterstock/prince of love
Generating speech with different rhythms and poses sounds more human-like, according to an evaluation of artificial intelligence trained on speeches taken from YouTube and podcasts.
Most artificial intelligence text-to-speech systems are trained on datasets of acted-out speech, so the output can sound lofty and one-dimensional. More natural speech often exhibits different rhythms and patterns to convey different meanings and emotions.
Now at Carnegie Mellon University in Pittsburgh, Pennsylvania, Alexander Rudnicky…