Language models struggle to create kinship terms
Summary and headline written by AI from the source article. How we work
Researchers tested five large language models. GPT OSS120B, Llama 3.370B, GLM-5.1, and two others, by asking them to generate kinship terms in Hindi, Tamil, and Korean.
The team compared how well the models recognised the correct terms (from a multiple-choice list) with how well they created them from scratch. The models excelled at choosing the right answer when given options, with GPT OSS120B correctly selecting terms 90.67% of the time.
However, when asked to produce the terms themselves, GPT OSS120B only succeeded 36% of the time. The study suggests that current tests may not accurately measure a language model’s understanding of cultural concepts. The researchers found that performance varied significantly depending on the specific language and the way the question was phrased.
They noted a “paternal-lineage advantage” in Hindi, meaning the models were better at identifying relationships through the father’s side of the family, but this pattern did not hold true in Korean. The team used Tamil as a control language to account for variations in measurement. They argue that evaluating language models through generation tasks, asking them to create language, is crucial alongside traditional multiple-choice tests.


