Graduation Year

2026

Document Type

Dissertation

Degree

Ph.D.

Degree Name

Doctor of Philosophy (Ph.D.)

Degree Granting Department

Mathematics and Statistics

Major Professor

Kandethody M. Ramachandran, Ph.D.

Committee Member

Lu Lu, Ph.D.

Committee Member

Christos P. Tsokos, Ph.D.

Committee Member

Tapas K. Das, Ph.D.

Keywords

deep learning, Deep reinforcement learning, offline reinforcement learning, online reinforcement learning, portfolio optimization, risk measure

Abstract

Portfolio optimization is important in investment, with the purpose of maximizing profit whilemanaging risk. The traditional portfolio optimization method, Modern Portfolio Theory, is based on the assumption that returns are normally distributed, which is always violated in practical financial market settings. In addition, traditional approaches to portfolio optimization are mostly designed as static optimizations and fail to account for the nonlinearity and non-stationarity of contemporary financial markets. Recent developments in deep reinforcement learning have attracted considerable attention for treating portfolio optimization as a sequential decision-making process. Traditional portfolio optimization methods usually evaluate risk based on the variance of the returns. Deep reinforcement learning methods can integrate the entropic risk measure, which accounts for uncertainty, into the action selection mechanism. We developed a risk-aware deep reinforcement learning method by integrating an entropic risk measure into the actor and critic networks of Deep Deterministic Policy Gradient to enhance the robustness of portfolio optimization under uncertainty. Furthermore, we extend the proposed approach by adding episodic memory to the deep reinforcement learning paradigm. Episodic memory stores past information and utilizes high-reward experiences during learning. However, online exploration in deep reinforcement learning can be risky due to the possibility of suffering financial losses caused by poor exploratory actions. Although offline reinforcement learning methods rely on historical data and avoid interaction with the environment, they suffer from distributional shift and exhibit poor online adaptation to environmental changes. We developed offline-online deep reinforcement learning methods that obtain a pre-trained policy from an offline dataset and adapt it to a new policy when an online dataset becomes available.

Share

COinS