ovr.news

Archaeology, rediscovered knowledge, the past opening up

Proverbs reveal multilingual Data’s power to unlock figurative language

arxiv.org · 24 September 2026

Summary and headline written by AI from the source article. How we work

Researchers analysed 742 proverb concepts, translated into 6,787 instances across seven languages, to determine how multilingual training data impacts identifying figurative language.

The team developed a new way to categorise proverbs, looking at four types of figurative meaning: metaphorical language, moral or advisory statements, cause-and-effect relationships, and culturally specific expressions. They tested five different language models, including large language models adjusted with specific instructions.

The research demonstrates that using around half of the available translated multilingual data achieves almost the best possible performance in identifying figurative language. Combining these different types of figurative forms improved overall accuracy. Notably, proverbs relying on culture-specific meanings showed the greatest improvement when the models were trained with multilingual data. Instruction-tuned large language models benefited most from learning to identify moral/advisory and culture-specific proverbs.

This suggests that focusing on a broader range of figurative language, beyond just metaphors, improves a model’s ability to understand proverbs. The findings encourage researchers to develop more complex frameworks for analysing figurative language that account for these varied meanings at the level of underlying concepts.

Was this worth your time?
Read on arxiv.org
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this

More in Discovery

Browse all Discovery articles

What made it worth it?

What's wrong with this article?