Code-Guided Reasoning in Vision-LanguageModels for Complex Diagram Understanding
Understanding complex structured diagrams, such as circuit schematics, molecular structures, musical notation, or business process models, requires precise symbolic, spatial, and relational reasoning. Current vision-language models (VLMs) struggle with such tasks because theylack access to the underlying symbolic structure that governs these diagrams. We introduce a training paradigm in which VLMs explicitly learn to reason through an intermediate symbolic representation of the image that is expressed in code. We generate a large synthetic dataset covering 21 diagram types across 7 domains by prompting large language models to generate code in specific formal representation languages (FRLs) and rendering them into paired code-image samples. During VLM training, the FRL code is provided along with the image, enabling the model to incorporate the symbolic representation during reasoning. Experiments show that models capable of producing valid code benefit from this symbolic intermediate layer, yielding improved accuracy on diagram understanding tasks. Our results demonstrate that integrating symbolic code into VLM training offers a promising direction for VLM design to handle complex visual data by bridging diagram perception with symbolic reasoning.
- Published in:
ESANN 2026 European Symposium on Artificial Neural Networks - Type:
Inproceedings - Authors:
- Year:
2026 - Source:
https://i6doc.com/en/book/?gcoi=28001100644460#h2tabFormats
Citation information
: Code-Guided Reasoning in Vision-LanguageModels for Complex Diagram Understanding, ESANN 2026 European Symposium on Artificial Neural Networks, 2026, August, i6doc, https://i6doc.com/en/book/?gcoi=28001100644460#h2tabFormats, Steinigen.etal.2026a,
@Inproceedings{Steinigen.etal.2026a,
author={Steinigen, Daniel; Flek, Lucie; Houben, Sebastian},
title={Code-Guided Reasoning in Vision-LanguageModels for Complex Diagram Understanding},
booktitle={ESANN 2026 European Symposium on Artificial Neural Networks},
month={August},
publisher={i6doc},
url={https://i6doc.com/en/book/?gcoi=28001100644460#h2tabFormats},
year={2026},
abstract={Understanding complex structured diagrams, such as circuit schematics, molecular structures, musical notation, or business process models, requires precise symbolic, spatial, and relational reasoning. Current vision-language models (VLMs) struggle with such tasks because theylack access to the underlying symbolic structure that governs these diagrams. We introduce a training paradigm in which...}}