← Back to projectsResearch

Identifiability and Explainability of Reinforcement Learning Models Using the Iowa Gambling Task

Python · SciPy · NumPy · emcee · metody Monte Carlo · statystyka bayesowska · modelowanie matematyczne · optymalizacja numeryczna

Modern artificial intelligence systems, while effective, often operate as 'black boxes'. In regulated industries (medicine, finance), it is crucial to use models that not only predict but also explain the decision-making mechanism.

As part of this project, I conducted a comprehensive stability and reliability audit of leading reinforcement learning (RL) models simulating decision-making behavior. I designed a novel architecture with significantly greater noise robustness and developed a rigorous, universal diagnostic pipeline that mathematically verifies when and to what extent the explanations generated by AI algorithms can be trusted.

 

Key achievements and engineering contributions

  • Design of a robust RL architecture (doubling model stability): I developed and tested a novel reinforcement learning model that combines asymmetric risk assessment (Prospect Theory) with dynamic uncertainty tracking (a Kalman Filter-inspired mechanism). The new architecture achieved a doubling of parameter estimation stability (increasing from ~38% in baseline models to as much as 76.2% of the entire working space), dramatically improving the reliability of generated explanations.
  • Building a simulation pipeline (stress-testing at 50k+ trials): I designed and implemented a multi-level validation pipeline based on multidimensional space sampling (Latin Hypercube Sampling) and bootstrap methods. This exposed critical flaws in the widely used point estimation approach (MLE) and demonstrated that standard error-reporting metrics (based on the Hessian matrix) exhibit extreme overconfidence, which in real-world deployments risks driving incorrect decisions based on random noise.
  • Implementation of a Bayesian model audit (overcoming optimizer limitations): I replaced classical, point-based gradient optimization with probabilistic Bayesian sampling using Monte Carlo methods (MCMC). I created a framework capable of extracting a clean information signal from highly noisy data sequences in 96% of cases where traditional gradient algorithms gave up and incorrectly classified parameters as entirely unrecoverable.
  • Discovery of a diagnostic dissociation (reliable classification despite noise): I developed a methodology for mapping the global behavioral space of decision-making systems (PSP) and overlaid it with parameter estimation error distributions. I demonstrated that high-level strategy classification and user segmentation remains fully stable and reliable, even when detailed quantitative parameters at the lowest level are subject to strong numerical noise. This enables the safe deployment of categorization systems with a very high degree of confidence.
  • Systematic bias risk mitigation in model personalization: I conducted a rigorous analysis of the impact of hierarchical parameter shrinkage methods. I demonstrated that popular averaging techniques mask unique, atypical behaviors (e.g., anomalies or rare pathologies) by artificially pulling them toward the 'norm'. I formulated a rigorous postulate for stress-testing models on 'clean data' (flat priors) to protect AI systems from algorithmic errors and unintended discrimination against unique profiles.

Let's talk.

Looking for a software engineer for your team or project? Get in touch.