Models struggle to read Ancient Greek without relying on guesswork
Summary and headline written by AI from the source article. How we work
Researchers compared how well vision-language models and traditional optical character recognition (OCR) systems read Ancient Greek texts.
They found that while both types of programs make mistakes, the vision-language models often produce grammatically correct but inaccurate interpretations, suggesting they rely more on their existing knowledge of the Greek language than on the images of the text itself. The study focused on critical editions of Ancient Greek, which are texts with scholarly notes and commentary.
Traditional OCR programs tended to produce random errors when they misread characters, while the vision-language models offered plausible substitutions. To understand this difference, the researchers altered the images of the texts and observed how each program responded. They discovered that when characters were changed, the traditional OCR remained more accurate, but the vision-language models quickly diverged from the correct text. The team also analyzed how much each model relied on the visual input during the reading process.
They found that a model specifically designed for OCR needed very little visual information to produce fluent text, even if it was incorrect. However, more general vision-language models still used the images, but sometimes made mistakes regardless. Attempts to correct the models during the reading process did not consistently improve accuracy, but fixing the text after it was generated did help.
The researchers suggest that evaluating these systems requires looking beyond overall accuracy to understand if the output is actually based on the visual evidence.