Leveraging Synthetically Generated Data for Real Estate Document Classification

Document classification in regulated domains like law, finance, or real estate is hindered by the scarcity of labeled data and strict privacy constraints. This paper presents a pipeline for synthetically generating training data for document classifiers using a combination of domain-specific templates, large language models, and data augmentation techniques. Focusing on two key document types relevant to real estate workflows, {\textless}em{\textgreater}Child Support Certificate and Refurbishment Roadmap{\textless}/em{\textgreater}, we construct realistic multi-page documents and generate negative classes using {LLM}-generated distractors. We train a {BERT}-based classifier on this synthetic dataset and evaluate it on real-world {OCR}-extracted documents, achieving strong performance despite the absence of real documents in training. Our findings highlight the feasibility of using synthetic data to overcome annotation bottlenecks and pave the way for broader applications in privacy-sensitive industries.

Citation information

Deußer, Tobias; Ramien, Gregor; Weber, Nico; Meidinger, Maximilian; Hahnbück, Max; Bauckhage, Christian; Sifa, Rafet: Leveraging Synthetically Generated Data for Real Estate Document Classification, 2025 IEEE International Conference on Big Data (BigData), 2025, December, {IEEE}, Institute of Electrical and Electronics Engineers, https://bonndoc.ulb.uni-bonn.de/xmlui/handle/20.500.11811/13972, Deusser.etal.2025c,

Associated Lamarr Researchers

Kopie von LAMARR Person 500x500 1 - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

Prof. Dr. Christian Bauckhage

Director to the profile
Prof. Dr. Rafet Sifa

Prof. Dr. Rafet Sifa

Principal Investigator Hybrid ML to the profile