NLP Colloquium with Dr. Jan Philip Wahle on “Can We Trust What Models Think and Say?”

As part of the NLP Colloquium series hosted by Lamarr NLP Area Chair and Principal Investigator Prof. Dr. Lucie Flek, we are pleased to welcome Dr. Jan Philip Wahle  from the University of Goettingen for a guest talk titled “Can We Trust What Models Think and Say?“. The talk will take place on September 09, 2026, from 11:00 to 12:00 p.m. (CET).

About the Talk

Can we trust a model simply because its outputs look correct, harmless, or well-reasoned? As language models become more capable, this question becomes increasingly difficult to answer. Undesirable behavior may remain hidden in seemingly benign outputs, while verbalized reasoning may not faithfully reflect the computations that drive a model’s decisions. In this talk, I will examine what it means to trust model behavior and how that trust can be measured. Using covert information leakage and reasoning faithfulness as two case studies, I will discuss where behavioral evaluation breaks down, what chain-of-thought can and cannot reveal, and how internal representations can help us distinguish safe-looking behavior from genuinely trustworthy model behavior.

About the Speaker

Dr. Jan Philip Wahle is a faculty member at the University of Göttingen. His research focuses on reasoning methods for large language models, including reinforcement learning and agentic systems, and on AI safety through mechanistic interpretability. He received his PhD in Computer Science from the University of Göttingen and was a visiting researcher at the National Research Council of Canada, working with Saif M. Mohammad. Before entering academia, he worked on machine learning for autonomous driving at Aptiv. His research has been published at venues including ACL and EMNLP and has received the ACL Best Resource Paper Award and the SemEval Best Task Award.

If you’d like to join, please subscribe to our mailing list here.
You’ll receive all details about the talk, including the Zoom link, via email.

Details

Date

9. September 2026

11:00 - 12:00

Location

University of Bonn

Friedrich-Hirzebruch-Allee 6, Room: 2.122

53115 Bonn

Topics

Natural Language Processing (NLP) , Education, General, Science

Tags

Event
Lamarr Events Colloquium - Lamarr Institute for Machine Learning (ML) and Artificial Intelligence (AI)

More events