TY - GEN
T1 - LaCoMSA: Language-Consistency Multilingual Self-Alignment with latent representation rewarding
AU - Tran, Khanh Tung
AU - O’Sullivan, Barry
AU - Nguyen, Hoang D.
N1 - © 2026, Association for Computational Linguistics.
PY - 2026/3/29
Y1 - 2026/3/29
N2 - Large Language Models (LLMs) have achieved impressive performance yet remain inconsistent across languages, often defaulting to high-resource outputs such as English. Existing multilingual alignment methods mitigate these issues through preference optimization but rely on external supervision, such as translation systems or English-biased signal. We propose Multilingual Self-Alignment (MSA), a targeted preference optimization framework that leverages an LLM’s own latent representations as intrinsic supervision signals, rewarding lower-resource language outputs based on their alignment with high-resource (English) counterparts in the “semantic hub”. We further introduce Language-Consistency MSA (LaCoMSA), which augments MSA with a final-layer language-consistency factor to prevent off-target generation. Integrated with Direct Preference Optimization, LaCoMSA improves a Llama 3 8B-based model multilingual win rates by up to 6.8% absolute (55.0% relatively) on X-AlpacaEval and achieves consistent gains across benchmarks and models. Our findings demonstrate that LaCoMSA can serve as an effective and scalable mechanism, opening a new venue toward multilingual self-alignment.
AB - Large Language Models (LLMs) have achieved impressive performance yet remain inconsistent across languages, often defaulting to high-resource outputs such as English. Existing multilingual alignment methods mitigate these issues through preference optimization but rely on external supervision, such as translation systems or English-biased signal. We propose Multilingual Self-Alignment (MSA), a targeted preference optimization framework that leverages an LLM’s own latent representations as intrinsic supervision signals, rewarding lower-resource language outputs based on their alignment with high-resource (English) counterparts in the “semantic hub”. We further introduce Language-Consistency MSA (LaCoMSA), which augments MSA with a final-layer language-consistency factor to prevent off-target generation. Integrated with Direct Preference Optimization, LaCoMSA improves a Llama 3 8B-based model multilingual win rates by up to 6.8% absolute (55.0% relatively) on X-AlpacaEval and achieves consistent gains across benchmarks and models. Our findings demonstrate that LaCoMSA can serve as an effective and scalable mechanism, opening a new venue toward multilingual self-alignment.
KW - Multilingual Self-Alignment
KW - Large Language Models (LLMs)
KW - [ComputerScience]
U2 - 10.18653/v1/2026.eacl-long.224
DO - 10.18653/v1/2026.eacl-long.224
M3 - Conference proceeding
AN - SCOPUS:105040580964
T3 - EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, Vol. 1 - (Long Papers)
SP - 4839
EP - 4853
BT - Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
A2 - Demberg, Vera
A2 - Inui, Kentaro
A2 - Marquez Villodre, Lluis
PB - Association for Computational Linguistics (ACL)
T2 - 19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026
Y2 - 24 March 2026 through 29 March 2026
ER -