Lecture: “Exploration vs. Exploitation in Reinforcement Learning”
Date & Time: October 16 (Friday), 2026, 3:40 ~ 5:30 PM
Place: Yang Doo Suk Hall, Kwanjeong Building, Seoul National University Library
In Reinforcement Learning, there is a traditional dilemma between actions being used for exploring the system, or for exploiting what is known about the system. This is also traditionally called the “Dual Control Problem” in Adaptive Control, where control can be used either to guide the system along its trajectory, or to excite the modes of the system to learn about it. We address the problem of how to optimally trade off between these two objectives.
Theoretically, the rate of growth of Regret can be established as optimal. Along with this, performance can be quantitatively studied on several examples of interest that have been studied through World Models and contemporary methods of interest.

Tutorial Lecture: The Reverse Biased Maximum Likelihood Estimate Method for Resolving the Exploration vs. Exploitation Dilemma
Date & Time: October 17 (Saturday), 2026, 9:30 AM ~ 12:30 PM
Place: Room B101, Building 43-2 dong, Seoul National University
The dilemma of Exploration vs. Exploitation can also be framed as an identifiability problem. When using a Certainty Equivalence approach, the parameter estimates converge to a limit, where under the control law corresponding to the limit, the true system is indistinguishable from the limiting value. This in turn introduces a certain bias in the parameter estimates that can be corrected for. This results in the Reverse Biased Maximum Likelihood Estimate (RBMLE) method, which produces optimal long-term average performance. It provides a graceful resolution of the Exploration vs. Exploitation dilemma.
• Handouts containing copies of the presentation slides will be provided.
2026 Lecturer: P.R. Kumar

P. R. Kumar received his B.Tech. from IIT Madras (1973) and D.Sc. from Washington University in St. Louis (1977). He was on the faculty at the University of Maryland, Baltimore County (1977–1984) and the University of Illinois at Urbana-Champaign (1985–2011). He has been at Texas A&M University since 2011.
His research spans machine learning, wireless networks, power systems, autonomous transportation, manufacturing systems, scheduling of wafer fabrication plants, adaptive control, game theory, and network information theory.
He is a member of the U.S. National Academy of Engineering, The World Academy of Sciences, and the Indian National Academy of Engineering. He is a Fellow of IEEE, ACM, and IFAC. His honors include an honorary doctorate from ETH Zurich, the IEEE Alexander Graham Bell Medal, the IEEE Control Systems Field Award, the ACM SIGMOBILE Outstanding Contribution Award, the Donald Eckman Award of AACC, the IEEE
Communication Society Ellersick Prize, the Infocom Achievement Award, the SIGMOBILE Test-of-Time Paper Award, and the Outstanding Contribution Award of COMSNETS.
He has received the Distinguished Alumnus Award of IIT Madras, the Alumni Achievement Award from Washington University in St. Louis, the Daniel C. Drucker Eminent Faculty Award of University of Illinois, and the University Distinguished Professor award from Texas A&M University.
