Fast rates for maximum entropy exploration
2023 · in (ICML 2023)
Abstract
We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The first type is visitation entropy maximization previously considered by Hazan et al. (2019) in the discounted setting. For this type of exploration, we propose a game-theoretic algorithm that has sample complexity thus improving the epsilon-dependence upon existing results. The second type of entropy we study is the trajectory entropy. This objective function is closely related to the entropy-regularized MDPs, and we propose a simple algorithm that has a sample complexity of order polynomial in S, A, H over epsilon. Interestingly, it is the first theoretical result in RL literature that establishes the potential statistical advantage of regularized MDPs for exploration. Finally, we apply developed regularization techniques to reduce sample complexity of visitation entropy maximization, yielding a statistical separation between maximum entropy exploration and reward-free exploration.
PDF · International Conference on Machine Learning · arXiv preprint · bibtex · DOI


