Zoom In Disparities in Healthcare {LLM} {Q\&A}
This paper systematically examines cross-lingual disparities in pre-training source and factuality alignment in Large Language Model (LLM) answers for multilingual healthcare Q\&A across English, German, Turkish, Chinese (Mandarin), and Italian. To support this analysis, we (i) constructed MultiWikiHealthCare, a multilingual dataset derived from Wikipedia; (ii) used it to examine cross-lingual differences in healthcare-related coverage; (iii) evaluated the alignment between LLM-generated responses and these reference sources; and (iv) conducted a case study on factual alignment through the use of contextual information and Retrieval-Augmented Generation (RAG). Our findings reveal substantial cross-lingual disparities in both Wikipedia coverage and LLM factual alignment. Based on our Wikipedia analysis, the lowest alignment is observed between English and Chinese Wikipedia pages. Across models, responses align more with English Wikipedia, even when the prompts are non-English. We further show that providing contextual excerpts from non-English Wikipedia at inference time effectively shifts factual alignment toward target knowledge.
- Published in:
Natural Language Processing and Information Systems - Type:
Inproceedings - Authors:
- Year:
2026 - Source:
https://link.springer.com/chapter/10.1007/978-3-032-29532-3_12
Citation information
: Zoom In Disparities in Healthcare {LLM} {Q\&A}, Natural Language Processing and Information Systems, 2026, 16696, 154--168, July, Springer Nature Switzerland, https://link.springer.com/chapter/10.1007/978-3-032-29532-3_12, Schlicht.etal.2026a,
@Inproceedings{Schlicht.etal.2026a,
author={Schlicht, Ipek Baris; Sayin, Burcu; Zhao, Zhixue; Labonté, Frederik M.; Barbera, Cesare; Viviani, Marco; Rosso, Paolo; Flek, Lucie},
title={Zoom In Disparities in Healthcare {LLM} {Q\&A}},
booktitle={Natural Language Processing and Information Systems},
volume={16696},
pages={154--168},
month={July},
publisher={Springer Nature Switzerland},
url={https://link.springer.com/chapter/10.1007/978-3-032-29532-3_12},
year={2026},
abstract={This paper systematically examines cross-lingual disparities in pre-training source and factuality alignment in Large Language Model (LLM) answers for multilingual healthcare Q\&A across English, German, Turkish, Chinese (Mandarin), and Italian. To support this analysis, we (i) constructed MultiWikiHealthCare, a multilingual dataset derived from Wikipedia; (ii) used it to examine...}}