AI safety and society
Balancing Global and Local Values in Multilingual AI
A multilingual alignment study reports lower harmful outputs without a measured drop in general response preferences across its test languages.

A safety system trained on one set of examples may miss harmful requests expressed in another language or cultural context. The Multilingual Alignment Prism study by Aakanksha and colleagues tests ways to reduce harmful outputs across languages while tracking how general responses change.
Training for local and global harms
The researchers use the 8-billion-parameter Aya 23 model and build a multilingual red-teaming dataset with human-annotated prompts. Their safety evaluation covers English, Hindi, French, Spanish, Russian and Arabic, separating “global” harm categories from harms that are more locally grounded. They compare supervised fine-tuning and Direct Preference Optimization (DPO), a method that trains a model from preferred and dispreferred responses. One DPO variant starts from a supervised fine-tuned checkpoint, called DPO(SFT) (Aakanksha et al., 2024, pp. 12027–12031).
On the study’s translated safety evaluation benchmark, SFT-Preferred reduced harmful generations by 56.6% and DPO(SFT) by 54.7% relative to the base model. For general open-ended responses, the variants had win rates of 67.4% and 71% against the base model, respectively (pp. 12031–12032). These metrics come from the paper’s evaluation design; they are not guarantees about every user, language or safety incident.
The researchers also report cross-harm transfer: training examples targeting global harms helped reduce local harmful outputs, and local examples helped with global harms. Their analysis found this pattern across the evaluated languages and training mixtures (pp. 12032–12034). Starting DPO from the supervised checkpoint performed better than starting from the instruction-tuned base model in their comparisons (pp. 12032–12033).
Safety remains a moving target
The study’s dataset covers a range of harm categories, but the authors acknowledge that it is incomplete. Some nuanced or context-specific harms may be absent, and harmful content changes over time. They describe the red-teaming data as spanning eight languages overall, while the principal safety evaluation is reported across six languages; neither set represents the full linguistic diversity of people who may use a model (pp. 12027, 12036).
The work tests one model family and selected training approaches. Its aggregate reductions can hide differences among prompts or languages, and general-response win rates measure judged outputs rather than long-term effects on users. The results are a controlled comparison of a training strategy, not a deployment audit.
Possible implications for practice
The findings suggest safety examples from one linguistic or cultural context may sometimes help in another. That offers a possible route to broader coverage when local examples are scarce. But the dataset limits show why transfer should be measured rather than assumed: a method that reduces one class of harmful responses may miss a less common local issue.
For organizations deploying multilingual systems, the study points toward evaluating safety and ordinary task quality together. A model that blocks harmful requests but loses usefulness can create a different failure. Local researchers and users may help identify what a benchmark overlooks, while repeated evaluation can track new forms of harm as they appear. The balance between shared safety rules and context-specific judgments remains a social and technical choice.
This paper evaluates one training configuration and a defined set of languages, so a deployment team would still need to test its own products, prompts and user populations. The results also illustrate why aggregate gains should be paired with language-level reporting: an overall reduction can obscure where safety improved, where general quality changed and where local harms remain underrepresented.
Bibliography & sources
- Aakanksha; Arash Ahmadian; Beyza Ermis; Seraphina Goldfarb-Tarrant; Julia Kreutzer; Marzieh Fadaee; Sara Hooker. 2024. “The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm.” In *Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, 12027–12049. Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.671.
