How Wrong Can Your Critic Be? Offline RL via the Evaluation Gap

Published:

Slides from my ICML 2026 talk. We rethink distribution shift in offline RL through the policy evaluation gap: choose policies whose value we can trust from the data, and turn that into a simple in-sample algorithm. Paper: REG: In-Sample RL via Regularizing the Evaluation Gap.