NLP Colloquium with Dr. Jan Philip Wahle on “Can We Trust What Models Think and Say?”
As part of the NLP Colloquium series hosted by Lamarr NLP Area Chair and Principal Investigator Prof. Dr. Lucie Flek, we are pleased to welcome Dr. Jan Philip Wahle from the University of Goettingen for a guest talk titled “Can We Trust What Models Think and Say?“. The talk will take place on September 09, 2026, from 11:00 to 12:00 p.m. (CET).
About the Talk
Can we trust a model simply because its outputs look correct, harmless, or well-reasoned? As language models become more capable, this question becomes increasingly difficult to answer. Undesirable behavior may remain hidden in seemingly benign outputs, while verbalized reasoning may not faithfully reflect the computations that drive a model’s decisions. In this talk, I will examine what it means to trust model behavior and how that trust can be measured. Using covert information leakage and reasoning faithfulness as two case studies, I will discuss where behavioral evaluation breaks down, what chain-of-thought can and cannot reveal, and how internal representations can help us distinguish safe-looking behavior from genuinely trustworthy model behavior.
About the Speaker
Dr. Jan Philip Wahle is a faculty member at the University of Göttingen. His research focuses on reasoning methods for large language models, including reinforcement learning and agentic systems, and on AI safety through mechanistic interpretability. He received his PhD in Computer Science from the University of Göttingen and was a visiting researcher at the National Research Council of Canada, working with Saif M. Mohammad. Before entering academia, he worked on machine learning for autonomous driving at Aptiv. His research has been published at venues including ACL and EMNLP and has received the ACL Best Resource Paper Award and the SemEval Best Task Award.
If you’d like to join, please subscribe to our mailing list here.
You’ll receive all details about the talk, including the Zoom link, via email.
Details
Date
9. September 2026
11:00 - 12:00
Location
University of Bonn
Friedrich-Hirzebruch-Allee 6, Room: 2.122
53115 Bonn
Topics
Natural Language Processing (NLP) , Education, General, Science