REG: In-Sample RL via Regularizing the Evaluation Gap
Published in ICML 2026, 2026
Offline reinforcement learning methods must avoid evaluating actions that are not supported by the data.
- We bound the off-policy evaluation gap and use Fenchel duality to turn the resulting optimization into an equivalent RL algorithm. This replaces IQL’s expectile regression with a simpler critic loss that comes with a theoretical guarantee on the gap.
- We propose an orthogonal policy gradient that adds on-policy information to the in-sample gradient, encouraging mode-seeking policies. It matches or outperforms a diffusion-policy baseline on D4RL MuJoCo and AntMaze.
- The learned critic can also be used to select good policies, reaching under 5% regret relative to the Top@10 trajectories.
