New metric assesses tone accuracy in multilingual speech synthesis
Summary and headline written by AI from the source article. How we work
Researchers developed DunDun, a new way to automatically evaluate how accurately text-to-speech systems pronounce tones in languages like Yor\`ub\'a, a Niger-Benue language spoken primarily in Nigeria.
Current automated metrics often fail because they disregard the pitch variations that define meaning in tonal languages. DunDun bypasses the need for labelled audio data by reading tone marks directly from the input text and comparing them to the pitch detected in the synthesized speech. The team validated DunDun in three ways.
They found that artificially flattening the pitch in speech samples lowered DunDun scores while leaving standard error rates unchanged. They also tested the metric by deliberately swapping high and low tones in a dataset of native recordings, confirming that DunDun consistently identified the errors. In a blind listening test, human listeners correctly identified tone-accurate speech clips 89.6% of the time, though the researchers note further study is needed to confirm if DunDun consistently aligns with these human judgements.
Applying DunDun to a multilingual text-to-speech model revealed that the model already produces relatively accurate tones in Yor\`ub\'a even before specific training for the language. Further training with just a few hours of audio data substantially reduced standard error rates while tone accuracy quickly plateaued. The researchers emphasize that the best metric for evaluating speech synthesis depends on the specific language, as existing metrics like error rate are sufficient for non-tonal languages like Swahili.
They have made the DunDun metric and validation process publicly available.


