ovr.news

Archaeology, rediscovered knowledge, the past opening up

English to Syriac translation advances endangered language tech

arxiv.org · 17 September 2026

Summary and headline written by AI from the source article. How we work

Researchers have created the first machine translation model for English to Assyrian Syriac, an endangered language spoken by an estimated 500,000 to 1,500,000 people worldwide.

Recognizing a gap in computational linguistics, the team built a dataset of 38,847 English-Syriac sentence pairs sourced from the Bible. This involved extracting text from PDF files, custom scripting for text segmentation, and manual review by three bilingual speakers to ensure accuracy.

The team used a phrase-based Statistical Machine Translation approach with the Moses framework, testing various model configurations to optimize performance. They preprocessed the Syriac text by removing diacritics, marks indicating pronunciation, and applying Byte-Pair Encoding to simplify the orthography of the complex Madnkhaya script. This script is used to write the East Syriac dialect of the language. The best model achieved a word-level BLEU score of 23.54, a standard metric for machine translation accuracy.

Further evaluation by 11 native Assyrian speakers yielded mean scores of 3.42 for adequacy and 3.34 for fluency, on a scale of 5. The researchers have made their dataset, scripts, and the trained model publicly available to support further research and language preservation efforts.

Was this worth your time?
Read on arxiv.org
Surfaced by the Discovery lens — one of the vital signs ovr.news reads.
How we evaluated this

More in Discovery

Browse all Discovery articles

What made it worth it?

What's wrong with this article?