ovr.news

Archaeology, rediscovered knowledge, the past opening up

Vietnamese speech corpus unites dialects and deepfake detection

arxiv.org · 25 September 2026

Summary and headline written by AI from the source article. How we work

Researchers created VietPrism, a large Vietnamese speech and deepfake corpus, to address limitations in current speech technology.

The collection includes 993.4 hours of real speech from 1,262 speakers and 3.1K hours of artificially generated speech. VietPrism uniquely combines transcripts, speaker identification, five Vietnamese dialect groups, and examples of code-switching, the mixing of Vietnamese and English within conversation, which makes up almost half of the corpus by length. The corpus allows for controlled testing of deepfake detection systems.

Initial evaluations of five multilingual detectors showed significant weaknesses, with error rates increasing as the artificial and real voices became more similar. Performance also varied considerably depending on the detector and the method used to create the fake speech. This suggests current detectors struggle with the nuances of Vietnamese speech and the challenge of identifying subtle differences between real and fabricated audio.

VietPrism offers a valuable resource for improving Vietnamese speech recognition and developing more reliable methods for detecting audio deepfakes. The researchers hope this will foster advancements in both fields, particularly given the increasing sophistication of audio manipulation technology.

Was this worth your time?
Read on arxiv.org
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this

More in Discovery

Browse all Discovery articles

What made it worth it?

What's wrong with this article?