A Three-Level Audit of LLM Alignment for Argument Quality Assessment

Large Language Models (LLMs) are increas- ingly used as automated evaluators of argument quality. However, existing studies typically as- sess models only through their agreement with human scores, leaving the reasoning process behind these judgments unexplored. In this pa- per, we propose a three-level audit framework for evaluating the reliability of LLM-based ar- gument quality assessment. The framework distinguishes between (1) surface alignment, measuring agreement between LLM-predicted scores and human annotations; (2) instruc- tional alignment, assessing whether generated rationales adhere to the intended evaluation cri- teria; and (3) faithfulness alignment, examin- ing whether predicted scores are supported by the generated rationales. To operationalize this audit, we introduce structural rationale prompt- ing,1 which guides LLMs to generate structured justifications before assigning scores across 11 dimensions of the Dagstuhl-15512 argument quality corpus. We evaluate several LLMs un- der this framework and find that structural ra- tionale prompting substantially improves agree- ment with human annotations compared to definition-based prompting. Furthermore, the generated rationales generally follow the evalu- ation instructions and remain highly consistent with the predicted scores. Overall, our results suggest that auditing LLM evaluators beyond surface score agreement provides deeper in- sight into the reliability and transparency of LLM-based argument quality assessment.

Citation information

Chen, Wei-Fan; Yu, Jinming; Flek, Lucie: A Three-Level Audit of LLM Alignment for Argument Quality Assessment, Proceedings of the 13th Workshop on Argument Mining andReasoning, 2026, 26--38, July, Association for Computational Linguistics, https://aclanthology.org/anthology-files/anthology-files/pdf/argmining/2026.argmining-1.pdf#page=35, Chen.etal.2026a,

Associated Lamarr Researchers

Prof. Dr. Lucie Flek

Prof. Dr. Lucie Flek

Area Chair NLP to the profile