Offline Reinforcement Learning

Published:

A tutorial on offline reinforcement learning: why distribution shift makes it hard, the main families of methods (policy constraints such as ReBRAC and DICE, in-sample methods such as IQL), and practical advice.