← Back to projectsResearch
When Interpretable Parameters Fail to Explain: Evidence from the Rescorla-Wagner Model
Python
Co-author of a scientific publication at ICCCI 2026 | Wrocław University of Science and Technology
Interpretable parameters in reinforcement learning (RL) models are commonly used to provide mechanistic explanations of human behavior. In this paper, we demonstrate that this practice can fail even for the minimal Rescorla-Wagner agent.
Key arguments and findings:
Anisotropic likelihood geometry: We show that vastly different combinations of learning rate and choice stochasticity can fit the same behavioral sequences (in a Probabilistic Learning Task) equally well under maximum likelihood estimation (MLE).
Estimation distortions: This phenomenon leads to structured errors — heavy-tailed scaled vector error (SVE) distributions, boundary hits, and overlapping recovered parameter distributions from different regimes. Grid-based diagnostics further reveal that parameter recovery quality varies substantially across the generating parameter space.
A novel diagnostic workflow: To expose these limitations, we propose a workflow that combines likelihood geometry analysis, boundary hit diagnostics, and grid-based parameter recovery verification.
Practical value: Our approach provides a procedure for assessing when fitted parameters are sufficiently constrained by the data to support reliable parameter-based explanations, and when such claims should be accompanied by appropriate caveats.
Let's talk.
Looking for a software engineer for your team or project? Get in touch.