AI trained on YouTube and podcasts speaks with ums and ahs

Digital waveform image

AI can generate more natural synthesized speech by including pauses

shutterstock/prince of love

Generating speech with different rhythms and poses sounds more human-like, according to an evaluation of artificial intelligence trained on speeches taken from YouTube and podcasts.

Most artificial intelligence text-to-speech systems are trained on datasets of acted-out speech, so the output can sound lofty and one-dimensional. More natural speech often exhibits different rhythms and patterns to convey different meanings and emotions.

Now at Carnegie Mellon University in Pittsburgh, Pennsylvania, Alexander Rudnicky…

Source link

Leave a Reply

Your email address will not be published. Required fields are marked *