Uploads from ‍김성범[ 교수 / 산업경영공학부 ]

Watch and track your favorite playlist.

Curated by: ‍김성범[ 교수 / 산업경영공학부 ] (327 videos)


Currently Playing: [Open DMQA Seminar] Outperforming Humans with Reinforcement Learning Agents in Atari Games

Arcade learning environment에서 Atari 게임은 심층 강화학습 에이전트의 성능을 측정하는 가장 대표적인 벤치마크이다. 이 벤치마크에서 사람보다 우수한 플레이를 보이는 에이전트를 생성하기 위해 오랫동안 다양한 연구들이 수행되었다. 그 결과 2020년에 International Conference on Machine Learning 학회에서 Agent57이라는 방법론이 발표가 되었고 Atari 게임 벤치마크가 제공하는 총 57개의 게임 전체에서 강화학습 에이전트가 사람보다 더 우수한 성능을 달성하는 결과를 보여주었다. 이번 세미나에서는 Agent57이 만들어지기까지 밑바탕이 되었던 연구들을 시간 순서대로 하나씩 소개하고자 한다. [1] Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., ... & Hassabis, D. (2015). Human-level control through deep reinforcement learning. nature, 518(7540), 529-533. [2] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2015). Prioritized experience replay. arXiv preprint arXiv:1511.05952. [3] Nair, A., Srinivasan, P., Blackwell, S., Alcicek, C., Fearon, R., De Maria, A., ... & Silver, D. (2015). Massively parallel methods for deep reinforcement learning. arXiv preprint arXiv:1507.04296. [4] Horgan, D., Quan, J., Budden, D., Barth-Maron, G., Hessel, M., van Hasselt, H., & Silver, D. (2018, February). Distributed Prioritized Experience Replay. In International Conference on Learning Representations. [5] Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., & Dabney, W. (2018, September). Recurrent experience replay in distributed reinforcement learning. In International conference on learning representations. [6] Badia, A. P., Sprechmann, P., Vitvitskyi, A., Guo, D., Piot, B., Kapturowski, S., ... & Blundell, C. Never Give Up: Learning Directed Exploration Strategies. In International Conference on Learning Representations. [7] Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z. D., & Blundell, C. (2020, November). Agent57: Outperforming the atari human benchmark. In International conference on machine learning (pp. 507-517). PMLR.


Tracks in this Playlist