76 episodes
- Dr Thomas Frost is an emergency physician based in London, UK. He is also in the final stages of completing a PhD at University College London, where he has been looking at offline reinforcement learning applied to healthcare settings.
Featured References
Robust Real-Time Mortality Prediction in the Intensive Care Unit using Temporal Difference Learning
Thomas Frost, Kezhi Li, Steve Harris — ML4H Symposium, PMLR 259, 2025
Insulin4RL: Real-Time Insulin Infusions for Offline Reinforcement Learning
Thomas Frost, Steve Harris — PhysioNet, 2026 (RRID:SCR_007345)
The Hidden Risks of Temporal Resampling in Clinical Reinforcement Learning
Thomas Frost, Hrisheekesh Vaidya, Steve Harris — arXiv preprint, 2026
Insulin4RL: Real-Time Insulin Management in the Intensive Care Unit for Offline Reinforcement Learning
Thomas Frost, Steve Harris — arXiv preprint, 2026
Additional References
The artificial intelligence clinician learns optimal treatment strategies for sepsis in intensive care — Komorowski et al. 2018
Off by a beat: the effects of temporal misalignment in reinforcement learning for sepsis treatment — Tang et al. 2026
Identifying Decision Points for Safe and Interpretable Reinforcement Learning in Hypotension Treatment — Zhang et al. 2021
Where do doctors disagree? Characterizing Decision Points for Safe Reinforcement Learning in Choosing Vasopressor Treatment — Brown et al. 2025
Loss of plasticity in deep continual learning — Dohare et al. 2024 - Joseph Modayil is the Founder, President & Research Director of Openmind Research Institute.
Featured ReferencesÂ
Openmind Research InstituteÂ
The Alberta Plan for AI ResearchÂ
Richard S. Sutton, Michael Bowling, Patrick M. PilarskiÂ
Additional References Â
Joseph Modayil on Google Scholar Â
Joseph Modayil Homepage - Danijar Hafner was a Research Scientist at Google DeepMind until recently.
Featured References Â
Training Agents Inside of Scalable World Models [ blog ]Â
Danijar Hafner, Wilson Yan, Timothy Lillicrap
One Step Diffusion via Shortcut Models
Kevin Frans, Danijar Hafner, Sergey Levine, Pieter Abbeel
Action and Perception as Divergence Minimization [ blog ]Â
Danijar Hafner, Pedro A. Ortega, Jimmy Ba, Thomas Parr, Karl Friston, Nicolas HeessÂ
Additional References Â
Mastering Diverse Domains through World Models [ blog ] DreaverV3l Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap Â
Mastering Atari with Discrete World Models [ blog ] DreaverV2 ; Danijar Hafner, Timothy Lillicrap, Mohammad Norouzi, Jimmy Ba Â
Dream to Control: Learning Behaviors by Latent Imagination [ blog ] Dreamer ; Danijar Hafner, Timothy Lillicrap, Jimmy Ba, Mohammad NorouziÂ
Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos [ Blog Post ], Baker et al - David Abel is a Senior Research Scientist at DeepMind on the Agency team, and an Honorary Fellow at the University of Edinburgh. His research blends computer science and philosophy, exploring foundational questions about reinforcement learning, definitions, and the nature of agency. Â
Featured References Â
Plasticity as the Mirror of Empowerment Â
David Abel, Michael Bowling, André Barreto, Will Dabney, Shi Dong, Steven Hansen, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland, Tom Schaul, Satinder Singh Â
A Definition of Continual RL Â
David Abel, André Barreto, Benjamin Van Roy, Doina Precup, Hado van Hasselt, Satinder Singh Â
Agency is Frame-Dependent Â
David Abel, André Barreto, Michael Bowling, Will Dabney, Shi Dong, Steven Hansen, Anna Harutyunyan, Khimya Khetarpal, Clare Lyle, Razvan Pascanu, Georgios Piliouras, Doina Precup, Jonathan Richens, Mark Rowland, Tom Schaul, Satinder Singh Â
On the Expressivity of Markov Reward Â
David Abel, Will Dabney, Anna Harutyunyan, Mark Ho, Michael Littman, Doina Precup, Satinder Singh — Outstanding Paper Award, NeurIPS 2021 Â
Additional References Â
Bidirectional Communication Theory — Marko 1973 Â
Causality, Feedback and Directed Information — Massey 1990 Â
The Big World Hypothesis — Javed et al. 2024 Â
Loss of plasticity in deep continual learning — Dohare et al. 2024 Â
Three Dogmas of Reinforcement Learning — Abel 2024 Â
Explaining dopamine through prediction errors and beyond — Gershman et al. 2024 Â
David Abel Google Scholar Â
David Abel personal website Jake Beck, Alex Goldie, & Cornelius Braun on Sutton's OaK, Metalearning, LLMs, Squirrels @ RLC 2025
2025/08/19 | 12 mins.Recorded at Reinforcement Learning Conference 2025 at University of Alberta, Edmonton Alberta Canada.
Featured References
Lecture on the Oak Architecture, Rich Sutton
Alberta Plan, Rich Sutton with Mike Bowling and Patrick Pilarski
Additional References
Jacob Beck on Google ScholarÂ
Alex Goldie on Google Scholar
Cornelius Braun on Google Scholar
Reinforcement Learning Conference
More Technology podcasts
Trending Technology podcasts
About TalkRL: The Reinforcement Learning Podcast
TalkRL podcast is All Reinforcement Learning, All the Time.
In-depth interviews with brilliant people at the forefront of RL research and practice.
Guests from places like MILA, OpenAI, MIT, DeepMind, Berkeley, Amii, Oxford, Google Research, Brown, Waymo, Caltech, and Vector Institute.
Hosted by Robin Ranjit Singh Chauhan.
Podcast websiteListen to TalkRL: The Reinforcement Learning Podcast, Waveform: The MKBHD Podcast and many other podcasts from around the world with the radio.net app

Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features
Get the free radio.net app
- Stations and podcasts to bookmark
- Stream via Wi-Fi or Bluetooth
- Supports Carplay & Android Auto
- Many other app features


TalkRL: The Reinforcement Learning Podcast
Scan code,
download the app,
start listening.
download the app,
start listening.

































