Robust portfolio optimization using tail-risk-aware fuzzy distributional reinforcement learning

Document Type : Original Manuscript

Authors

College of Management and Economics, Tianjin University, Tianjin 300072, China

Abstract

This research presents the Tail-Risk-Aware Fuzzy Distributional Reinforcement Learning (T-FDRL) framework, an innovative method for robust portfolio management in dynamic, non-stationary financial markets. By combining Fuzzy Markov Decision Processes, distributional reinforcement learning, and Wasserstein-based robust optimization, T-FDRL tackles uncertainty in market signals, heavy-tailed return distributions, and shifts in market dynamics. It employs fuzzy sets to model vague market data, captures the complete return distribution, and optimizes portfolio allocations to maximize returns while minimizing tail risks using Conditional Value at Risk (CVaR). Tested on historical data from Apple, Amazon, Google, and Microsoft stocks spanning 2016 to 2025, T-FDRL delivers a cumulative return of 1544.21%, an annualized return of 21.22%, and a Sharpe Ratio of 1.60, surpassing benchmarks like Deep Q-Networks (272.38%), Fuzzy Reinforcement Learning (1120.22%), Value-Based DQN (1067.86%), and Robust Markov Decision Processes (652.63%). With low volatility (13.23%) and a CVaR 5% of -1.64%, T-FDRL excels in risk management, especially during turbulent periods such as 2020 and 2022. This framework provides a resilient, adaptive approach to portfolio optimization, offering valuable insights for investors and policymakers navigating complex financial landscapes.

Keywords

Main Subjects


[1] P. Artzner, F. Delbaen, J. M. Eber, D. Heath, Coherent measures of risk, Mathematical Finance, 9(3) (2001), 203-228. https://doi.org/10.1111/1467-9965.00068
[2] R. Ayari, Feature configuration effects in DRL portfolio management: A risk-focused evaluation under market stress, Quantitative Finance, 26(1) (2026), 119-136. https://doi.org/10.1080/14697688.2025.2592822
[3] T. Bauman, L. Mrčela, S. Goluža, Z. Kostanjčar, A deep learning approach to goal-based portfolio optimization in non-stationary environments, IEEE Access, 13 (2025), 128158-128172. https://doi.org/10.1109/ACCESS.2025. 3588247
[4] S. D. Bekiros, Heterogeneous trading strategies with adaptive fuzzy actor-critic reinforcement learning: A behavioral approach, Journal of Economic Dynamics and Control, 34(6) (2010), 1153-1170. https://doi.org/10.1016/j. jedc.2010.01.015
[5] R. Bellman, A Markovian decision process, Journal of Mathematics and Mechanics, 6(5) (1957), 679-684. https: //www.jstor.org/stable/24900506
[6] I. S. Benistan, M. J. Shahbazzadeh, M. Eslami, Addressing lightning and market uncertainties in self-scheduling: A fuzzy-Markov approach for smart grids, Scientific Reports, 16 (2026), 8923. https://doi.org/10.1038/ s41598-026-42588-8
[7] N. P. Bhatta, F. Amsaad, Ml assisted techniques in power side channel analysis for trojan classification, Cluster Computing, 28(3) (2025), 157. https://doi.org/10.1007/s10586-024-04715-w
[8] A. Birashk, L. Khan, Federated continual learning for task-incremental and class-incremental problems: A survey, Expert Systems with Applications, 297(Part C) (2025), 129278. https://doi.org/10.1016/j.eswa.2025.129278
[9] P. Bologna, A. Segura, Integrating stress tests within the Basel III capital framework: A macroprudentially coherent approach, Journal of Financial Regulation, 3(2) (2017), 159-186. https://doi.org/10.1093/jfr/fjx004
[10] H. Choudhary, A. Orra, K. Sahoo, M. Thakur, Risk-adjusted deep reinforcement learning for portfolio optimization: A multi-reward approach, International Journal of Computational Intelligence Systems, 18(1) (2025), 126. https: //doi.org/10.1007/s44196-025-00875-8
[11] H. Choudhary, A. Orra, M. Thakur, et al., A CVaR-constrained safe reinforcement learning framework with action repair for practical portfolio optimization, IEEE Transactions on Artificial Intelligence, 7 (2026), 1-15. https: //doi.org/10.1109/TAI.2026.3686783
[12] W. Dabney, M. Rowland, M. G. Bellemare, R. Munos, Distributional reinforcement learning with quantile regression, Proceedings of the AAAI Conference on Artificial Intelligence, 32(1) (2018), 2892-2901. https://doi.org/ 10.1609/aaai.v32i1.11791
[13] M. Ensafi, A. Mozdgir, N. G. Tehrani, O. Ahmadi, Development of a machine learning-based framework for improving appointment attendance prediction in outpatient clinics using stacking algorithm, International Journal of Industrial Engineering and Operational Research, 8(1) (2026), 68-94. https://doi.org/10.22034/ijieor.v8i1. 211
[14] M. Farhang, F. Safi-Esfahani, Recognizing mapreduce straggler tasks in big data infrastructures using artificial neu ral networks, Journal of Grid Computing, 18(4) (2020), 879-901. https://doi.org/10.1007/s10723-020-09514-2
[15] F. Ghorbani, D. Helic, D. Dan, Utilizing fuzzy reinforcement learning for prediction in volatile stock markets, Array, 30 (2026), 100741. https://doi.org/10.1016/j.array.2026.100741
[16] N. Gradojevic, R. Gençay, Fuzzy logic, trading uncertainty and technical trading, Journal of Banking and Finance, 37(2) (2013), 578-586. https://doi.org/10.1016/j.jbankfin.2012.09.012
[17] J. Gu, W. Du, X. Zhao, et al., Incorporating realistic margin constraints: A data-driven deep reinforcement learning framework for advanced portfolio management, IEEE Transactions on Knowledge and Data Engineering, (2026), 1 12. https://doi.org/10.1109/TKDE.2026.3701681
[18] Z. Hao, H. Zhang, Y. Zhang, Stock portfolio management by using fuzzy ensemble deep reinforcement learning algorithm, Journal of Risk and Financial Management, 16(3) (2023), 201. https://doi.org/10.3390/jrfm16030201
[19] T. Harnpadungkij, W. Chaisangmongkon, P. Phunchongharn, Risk-sensitive portfolio management by using distributional reinforcement learning, 2019 IEEE 10th International Conference on Awareness Science and Technology (iCAST), (2019), 1-6. https://doi.org/10.1109/ICAwST.2019.8923223
[20] Z. Hosseini-Nodeh, R. Khanjani-Shiraz, P. M. Pardalos, Portfolio optimization using robust mean absolute deviation model: Wasserstein metric approach, Finance Research Letters, 54 (2023), 103735. https://doi.org/10.1016/j. frl.2023.103735
[21] G. Hu, M. Gu, Markowitz meets Bellman: Knowledge-distilled reinforcement learning for portfolio management, arXiv preprint, (2024). https://doi.org/10.48550/arXiv.2405.05449
[22] M. Jalali, M. Najand, A. Cohen, Machine learning, thematic feature grouping, and the magnificent seven: A forecasting analysis, Journal of Risk and Financial Management, 19(4) (2026), 274. https://doi.org/10.3390/ jrfm19040274
[23] Z. Jiang, D. Xu, J. Liang, A deep reinforcement learning framework for the financial portfolio management problem, arXiv preprint, (2017). https://doi.org/10.48550/arXiv.1706.10059
[24] A. Z. Khan, P. Gupta, M. K. Mehlawat, A fuzzy rule-based system for portfolio selection using technical analysis, IEEE Transactions on Fuzzy Systems, 32(9) (2024), 4861-4875. https://doi.org/10.1109/TFUZZ.2024.3355515
[25] A. M. Khass, A. Cosse, V. Pandey, N. Motee, Conflict-aware active perception and control in 3D Gaussian splatting fields via control barrier functions, arXiv preprint, (2026). https://doi.org/10.48550/arXiv.2605.20566
[26] A. M. Khass, G. Liu, V. Pandey, et al., Active next-best-view optimization for risk-averse path planning, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2510.06481
[27] A. M. Khass, V. Pandey, G. Liu, et al., Multi-agent next-best-view optimization for risk-averse planning, arXiv preprint, (2026). https://arxiv.org/abs/2606.04158
[28] F. Khemlichi, H. Chougrad, Y. I. Khamlichi, et al., Deep deterministic policy gradient for portfolio management, 2020 6th IEEE Congress on Information Science and Technology (CiSt), (2021), 424-429. https://doi.org/10. 1109/CiSt49399.2021.9357266
[29] K. Kirtac, G. Germano, Leveraging LLM-based sentiment analysis for portfolio allocation with proximal policy op timization, ICLR 2025 Workshop on Machine Learning Multiscale Processes, (2025). https://doi.org/10.18653/ v1/2025.realm-1.12
[30] D. Kuhn, P. M. Esfahani, V. A. Nguyen, S. Shafieezadeh-Abadeh, Wasserstein distributionally robust optimization: Theory and applications in machine learning, Operations Research and Management Science in the Age of Analytics, (2019), 130-166. https://doi.org/10.1287/educ.2019.0198
[31] K. Leballo, J. C. Mba, A parametric distributional reinforcement learning framework for conditional systemic risk estimation, International Journal of Data Science and Analytics, 22 (2026), 17. https://doi.org/10.1007/ s41060-025-00985-8
[32] J. Li, A deep reinforcement learning framework for financial portfolio management, arXiv preprint, (2017). https: //doi.org/10.48550/arXiv.2409.08426
[33] H. Li, M. Hai, Deep reinforcement learning model for stock portfolio management based on data fusion, Neural Processing Letters, 56(2) (2024). 108. https://doi.org/10.1007/s11063-024-11582-4
[34] Y. Li, P. Ni, V. Chang, Application of deep reinforcement learning in stock trading strategies and stock forecasting, Computing, 102(6) (2020), 1305-1322. https://doi.org/10.1007/s00607-019-00773-w
[35] Y. Li, W. Zheng, Z. Zheng, Deep robust reinforcement learning for practical algorithmic trading, IEEE Access, 7 (2019), 108014-108022. https://doi.org/10.1109/ACCESS.2019.2932789
[36] S. Liu, T. Cui, Y. Li, et al., AHRL-PM: Asynchronous hierarchical reinforcement learning framework for enhanced portfolio management, IEEE Transactions on Neural Networks and Learning Systems, (2026), 1-14. https://doi. org/10.1109/TNNLS.2026.3711337
[37] X. Y. Liu, H. Yang, J. Gao, C. D. Wang, FinRL: Deep reinforcement learning framework to automate trading in quantitative finance, Proceedings of the Second ACM International Conference on AI in Finance, (2021), 1-9. https://doi.org/10.1145/3490354.3494366
[38] S. Mashhadi, A. Saghezchi, V. G. Kashani, Interpretable machine learning for predicting startup funding, patenting, and exits, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2510.09465
[39] M. A. Masoudi, Robust deep reinforcement learning for portfolio management, PhD thesis, Universitèd’Ottawa/University of Ottawa, (2021). https://dx.doi.org/10.20381/ruor-26960
[40] R. Millar, J. Li, Bayesian optimization for CVaR-based portfolio optimization, Proceedings of the Genetic and Evolutionary Computation Conference, (2025), 1424-1432. https://doi.org/10.1145/3712256.3726307
[41] B. K. Mishra, M. Kumar, H. Fida, B. Kalaš, Risk-sensitive reinforcement learning for portfolio optimization under stochastic market dynamics, Mathematics, 14(8) (2026), 1334. https://doi.org/10.3390/math14081334
[42] S. Modaresi, V. Sachdeva, Y. Hu, et al., Evaluating zoning granularity in graph convolutional networks for predicting energy and structural performance, CumInCAD, (2025). https://doi.org/10.52842/conf.sigradi.2025.1.679
[43] A. Nasir, A. Khursheed, K. Ali, F. Mustafa, A Markov decision process model for optimal trade of options using sta tistical data, Computational Economics, 58(2) (2021), 327-346. https://doi.org/10.1007/s10614-020-10030-4
[44] R. L. Navaei, M. Safarzadeh, S. M. J. Sobhani, Optimizing flamelet generated manifold models: A machine learning performance study, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2507.01030
[45] P. Ndikum, S. Ndikum, Advancing investment frontiers: Industry-grade deep reinforcement learning for portfolio optimization, arXiv preprint, (2024). https://doi.org/10.48550/arXiv.2403.07916
[46] W. Nuipian, P. Meesad, M. Maliyaem, Innovative portfolio optimization using deep Q-network reinforcement learnng, Proceedings of the 2024 8th International Conference on Natural Language Processing and Information Retrieval, (2024), 292-297. https://doi.org/10.1145/3711542.3711567
[47] M. Omura, Y. Mukuta, K. Ota, et al., Offline reinforcement learning with Wasserstein regularization via optimal transport maps, arXiv preprint, (2025). https://arxiv.org/abs/2507.10843
[48] G. N. Raj, Adaptive and regime-aware RL for portfolio optimization, arXiv preprint, (2025). https://arxiv.org/ abs/2509.14385
[49] M. Rezaei, H. NezamAbAdipour, Integrating fuzzy logic with deep reinforcement learning to enhance financial portfolio management, Iranian Journal of Fuzzy Systems, 22(2) (2025), 187-204. https://doi.org/10.22111/ ijfs.2025.50135.8847
[50] A. Saghezchi, V. G. Kashani, F. Ghodratizadeh, A comprehensive optimization approach on financial resource allocation in Scale-Ups, Journal of Business and Management Studies, 6(6) (2024), 62-75. https://doi.org/10. 32996/jbms.2024.6.6.5
[51] J. Schulman, F. Wolski, P. Dhariwal, et al., Proximal policy optimization algorithms, arXiv preprint, (2017). https://doi.org/10.48550/arXiv.1707.06347
[52] A. Sharma, F. Chen, J. Noh, et al., Hedging beyond the mean: A distributional reinforcement learning perspective for hedging portfolios with structured products, arXiv preprint, (2024). https://arxiv.org/abs/2407.10903
[53] T. Skeepers, T. L. Van Zyl, A. Paskaramoorthy, MA-FDRNN: Multi-asset fuzzy deep recurrent neural network reinforcement learning for portfolio management, 2021 8th International Conference on Soft Computing and Machine Intelligence (ISCMI), (2021), 32-37. https://doi.org/10.1109/ISCMI53840.2021.9654987
[54] R. Venugopal, C. Veeramani, S. A. Edalatpanah, Enhancing daily stock trading with a novel fuzzy indicator: Performance analysis using Z-number based fuzzy TOPSIS method, Results in Control and Optimization, 14 (2024), 100365. https://doi.org/10.1016/j.rico.2023.100365
[55] X. Wang, L. Liu, Risk-sensitive deep reinforcement learning for portfolio optimization, Journal of Risk and Financial Management, 18(7) (2025), 347. https://doi.org/10.3390/jrfm18070347
[56] W. Wiesemann, D. Kuhn, B. Rustem, Robust Markov decision processes, Mathematics of Operations Research, 38(1) (2013), 153-183. https://doi.org/10.1287/moor.1120.0566
[57] H. Yang, X. Y. Liu, S. Zhong, A. Walid, Deep reinforcement learning for automated stock trading: An ensemble strategy, Proceedings of the First ACM International Conference on AI in Finance, (2020), 1-8. https://doi.org/ 10.1145/3383455.3422540
[58] Y. Yang, T. Wang, Y. Fu, et al., Portfolio management based on value distribution reinforcement learning algorithm, Frontiers in Artificial Intelligence, 8 (2026), 1709493. https://doi.org/10.3389/frai.2026.1709493
[59] M. Younesi Heravi, I. Jeong, Y. Jang, A vision-based approach for human activity intensity estimation using kinematic features, Journal of Computing in Civil Engineering, 40(6) (2026), 04026092. https://doi.org/10. 1061/JCCEE5.CPENG-7683
[60] L. A. Zadeh, Fuzzy sets, Information and Control, 8(3) (1965), 338-353. https://doi.org/10.1016/ S0019-9958(65)90241-X
[61] Y. Zhang, X. Li, S. Guo, Portfolio selection problems with Markowitz’s mean-variance framework: A re view of literature, Fuzzy Optimization and Decision Making, 17(2) (2018), 125-158. https://doi.org/10.1007/ s10700-017-9266-z
[62] Y. T. Zhang, J. Y. Yang, Y. Wu, Dynamic resource allocation strategy of multi-objective fuzzy optimization based on Markov decision process, IEEE Access, 11 (2023), 99607-99613. https://doi.org/10.1109/ACCESS.2023.3314657
[63] Z. Zhang, S. Zohren, S. Roberts, Deep reinforcement learning for trading, arXiv preprint, (2020). https://doi. org/10.48550/arXiv.1911.10107