Textual data bias detection and mitigation – an extensible pipeline with experimental evaluation

Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations, encompassing harmful language and skewed demographic distributions. Regulations such as the EU AI Act require identifying and mitigating biases against protected groups in data, with the ultimate goal of preventing unfair model outputs. However, practical guidance and operationalization are lacking. We propose a comprehensive data bias detection and mitigation pipeline comprising four components that address two data bias types, namely representation bias and (explicit) stereotypes for a configurable sensitive attribute. First, we leverage LLM-generated word lists, created according to defined quality criteria, to detect relevant group labels. Second, representation bias is quantified using the Demographic Representation Score. Third, we detect and mitigate stereotypes using sociolinguistically informed filtering. Finally, we mitigate representation bias through Grammar- and Context-Aware Counterfactual Data Augmentation. We conduct a twofold evaluation using gender, religion, and age as examples. First, we evaluate the effectiveness of each individual component on data debiasing through human validation and baseline comparison. The findings demonstrate that we successfully reduce representation bias and (explicit) stereotypes in a text dataset. Second, we evaluate the effect of data debiasing on model bias by benchmarking several models (0.6B-8B parameters) fine-tuned on the debiased text dataset. This evaluation reveals that LLMs fine-tuned on debiased data do not consistently show improved performance on bias benchmarks. These results expose critical gaps in current evaluation methodologies and highlight the need for targeted data interventions to address manifested model bias.

  • Published in:
    Expert Systems with Applications
  • Type:
    Article
  • Authors:
    Görge, Rebekka; Gannamaneni, Sujan Sai; Naeven, Tabea; Abdelwahab, Hammam; Allende-Cid, Héctor; Cremers, Armin B.; Helmer, Lennard; Mock, Michael; Schmitz, Anna; Xue, Songkai; Yildirir, Elif; Poretschkin, Maximilian; Wrobel, Stefan
  • Year:
    2026

Citation information

Görge, Rebekka; Gannamaneni, Sujan Sai; Naeven, Tabea; Abdelwahab, Hammam; Allende-Cid, Héctor; Cremers, Armin B.; Helmer, Lennard; Mock, Michael; Schmitz, Anna; Xue, Songkai; Yildirir, Elif; Poretschkin, Maximilian; Wrobel, Stefan: Textual data bias detection and mitigation – an extensible pipeline with experimental evaluation, Expert Systems with Applications, 2026, 330, 132995, December, Elsevier BV, Goerge.etal.2026a,

Associated Lamarr Researchers

lamarr institute person Gorge Rebekka - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

Rebekka Görge

Autorin to the profile
lamarr institute person Gannamaneni Sujan Sai e1663925008286 - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

Sujan Sai Gannamaneni

Author to the profile
lamarr institute person Mock Michael - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

Dr. Michael Mock

Author to the profile
lamarr institute person Poretschkin - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

Dr. Maximilian Poretschkin

Autor to the profile
lamarr institute person Wrobel Stefan e1663925461852 - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

Prof. Dr. Stefan Wrobel

Director to the profile