Language Models reflect cultural values differently
Summary and headline written by AI from the source article. How we work
Researchers built a dataset to examine whether large language models apply the same values when processing different languages.
The team focused on Chinese Social Values, a system of beliefs encompassing national, societal, and personal ethics. They created C-Voices, a dataset of 86,400 scenarios in six languages, each presenting a dilemma with one action aligning with these values and another conflicting with them. The researchers then developed a method to influence the models’ responses without retraining them.
This method identifies how a model’s internal workings change when presented with value-aligned versus conflicting choices, then subtly adjusts the model’s processing to favour the desired value. Experiments across six languages revealed that language models don’t consistently prioritize the same values. A single dilemma could produce different responses depending on the language used.
The team’s steering method successfully guided the models toward Chinese Social Values and demonstrated the ability to transfer these value preferences between languages, working with existing value-alignment tools like FLAMES and ValuePrism.


