Graduation Year
2026
Document Type
Dissertation
Degree
Ph.D.
Degree Name
Doctor of Philosophy (Ph.D.)
Degree Granting Department
Mathematics and Statistics
Major Professor
Kandethody M. Ramachandran, Ph.D.
Committee Member
Lu Lu, Ph.D.
Committee Member
Christos P. Tsokos, Ph.D.
Committee Member
Tapas K. Das, Ph.D.
Keywords
deep learning, Deep reinforcement learning, offline reinforcement learning, online reinforcement learning, portfolio optimization, risk measure
Abstract
Portfolio optimization is important in investment, with the purpose of maximizing profit whilemanaging risk. The traditional portfolio optimization method, Modern Portfolio Theory, is based on the assumption that returns are normally distributed, which is always violated in practical financial market settings. In addition, traditional approaches to portfolio optimization are mostly designed as static optimizations and fail to account for the nonlinearity and non-stationarity of contemporary financial markets. Recent developments in deep reinforcement learning have attracted considerable attention for treating portfolio optimization as a sequential decision-making process. Traditional portfolio optimization methods usually evaluate risk based on the variance of the returns. Deep reinforcement learning methods can integrate the entropic risk measure, which accounts for uncertainty, into the action selection mechanism. We developed a risk-aware deep reinforcement learning method by integrating an entropic risk measure into the actor and critic networks of Deep Deterministic Policy Gradient to enhance the robustness of portfolio optimization under uncertainty. Furthermore, we extend the proposed approach by adding episodic memory to the deep reinforcement learning paradigm. Episodic memory stores past information and utilizes high-reward experiences during learning. However, online exploration in deep reinforcement learning can be risky due to the possibility of suffering financial losses caused by poor exploratory actions. Although offline reinforcement learning methods rely on historical data and avoid interaction with the environment, they suffer from distributional shift and exhibit poor online adaptation to environmental changes. We developed offline-online deep reinforcement learning methods that obtain a pre-trained policy from an offline dataset and adapt it to a new policy when an online dataset becomes available.
Scholar Commons Citation
Zhang, Shimin, "Risk-Aware Deep Reinforcement Learning in Portfolio Optimization" (2026). USF Tampa Graduate Theses and Dissertations.
https://digitalcommons.usf.edu/etd/11455
