Fuzzy logic in large language models: Concepts, applications, and challenges

Document Type : Original Manuscript

Authors

1 Department of Computer Engineering, Faculty of Engineering, Lorestan University, Khorramabad, Iran

2 Department of Computer Engineering, Technical and Vocational University (TVU), Tehran, Iran.

10.22111/ijfs.2026.10100

Abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation; however, they continue to suffer from critical limitations, including predictive uncertainty, factual hallucination, and lack of interpretability. Fuzzy logic, with its inherent capacity to model graded membership and reason under imprecision, has emerged as a promising complementary paradigm to mitigate these shortcomings. Despite the rapidly growing body of literature, a systematic review that comprehensively synthesizes integration methodologies, empirical gains, and deployment challenges in fuzzy-LLM systems remains absent. This study conducts a systematic literature review of 70 peer-reviewed studies published between November 2022 and June 2026, organized around five research questions: (1) theoretical frameworks for fuzzy-LLM integration, (2) improvements in reasoning performance, (3)  allucination mitigation mechanisms, (4) fuzzy evaluation metrics, and (5) practical implementation obstacles. Our analysis reveals that fuzzy rule-based systems dominate the integration landscape, comprising 40 studies (57.1%), whereas fuzzy attention mechanisms remain critically underexplored, with only 3 studies (4.3%). Uncertainty reasoning constitutes the predominant application domain, addressed in 54 studies (77.1%), while causal reasoning remains entirely uninvestigated. Empirical findings indicate that fuzzy mechanisms yield measurable accuracy improvements between 1.5% and 3.0%, alongside substantial hallucination reduction. Among mitigation strategies, fuzzy refinement emerges as both the most  frequently adopted and effective approach, reported in 19 studies (27.1%). Nevertheless, large-scale deployment continues to face persistent barriers, particularly computational overhead and rule explosion, partially mitigated through  Type-2 fuzzy systems and differentiable fuzzy layers. This review identifies fuzzy attention, standardized benchmarks, and  causal reasoning as pivotal directions for future inquiry, offering a foundational reference for researchers and  practitioners developing reliable, interpretable, and uncertainty-aware LLM-based systems. 

Keywords


[1] N. Akbari, et al., CognitiveContinuum: Intent-aware orchestration using LLM and fuzzy logic, (2026). https: //nzjohng.github.io/publications/papers/icws2026.pdf
[2] H. M. Alabool, Large language model evaluation criteria framework in healthcare: Fuzzy MCDM approach, SN Computer Science, 6(1) (2025), 57. https://doi.org/10.1007/s42979-024-03533-6
[3] A. H. Alamoodi, et al., A novel evaluation framework for medical LLMs: Combining fuzzy logic and MCDM for medical relation and clinical concept extraction, Journal of Medical Systems, 48(1) (2024), 81. https://doi.org/ 10.1007/s10916-024-02090-y
[4] A. E. Amer, M. Amer, Using multi-agent architecture to mitigate the risk of LLM hallucinations, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2507.01446
[5] F. Anhao, et al., Integrating large language models into a novel intuitionistic fuzzy PROBID method for multi criteria decision-making problems, Mathematics, 13(17) (2025), 2878. https://doi.org/10.3390/math13172878
[6] Z. Anwar, et al., Fuzzy ensemble of fined tuned BERT models for domain-specific sentiment analysis of software engineering dataset, Plos One, 19(5) (2024). https://doi.org/10.1371/journal.pone.0300279
[7] M. L. Bangerter, et al., A hybrid framework integrating llm and anfis for explainable fact-checking, IEEE Transactions on Fuzzy Systems, 33(12) (2024). https://doi.org/10.1109/TFUZZ.2024.3431710
[8] T. B. Brown, et al., Language models are few-shot learners, in Advances in Neural Information Processing Systems (NeurIPS), (2020), 1877-1901.
[9] J. P. Carvalho, R. Ribeiro, Fuzzy fingerprints in limited discrete feature spaces, International Conference on Information Processing and Management of Uncertainty in Knowledge-Based Systems, Springer, Cham, (2026). https://doi.org/10.1007/978-3-032-29000-7_21
[10] S. Chakraborty, F. Heintz, Enhancing time series forecasting with fuzzy attention-integrated transformers, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2504.00070
[11] H. Chen, et al., Enhancing educational Q&A systems using a Chaotic Fuzzy Logic-Augmented large language model, Frontiers in Artificial Intelligence, 7 (2024), 1404940. https://doi.org/10.3389/frai.2024.1404940
[12] P. Chen, et al., Fuzzy reasoning chain (FRC): An innovative reasoning framework from fuzziness to clarity, Findings of the Association for Computational Linguistics: EMNLP, (2025), 10230-10240. https://doi.org/10.18653/v1/ 2025.findings-emnlp.541
[13] A. Chikhalikar, et al., Open vocabulary object search utilizing large language models and fuzzy inferencing, 2025 IEEE/SICE International Symposium on System Integration (SII), (2025). https://doi.org/10.1109/SII59315. 2025.10870891
[14] D. Dimitrov, Towards more reliable SQL auto-grading: A hybrid approach Using LLMs, intuitionistic fuzzy sets, and traditional methods, International Conference on Flexible Query Answering Systems, (2025), 253-264. https: //doi.org/10.1007/978-3-032-05607-8_24
[15] Y. Dong, T. Ito, A linguistic negotiation for consensus-making among LLM agents with fuzzy utility functions, IEICE Transactions on Information and Systems, E109.D(7) (2026), 1112-1122. https://doi.org/10.1587/ transinf.2025EDP7179
[16] M. B. Dowlatshahi, S. Beiranvand, Optimizing deep Q-networks with fuzzy inference-based adaptive replay buffer management, Iranian Journal of Fuzzy Systems, 22(4) (2025), 161-174. https://doi.org/10.22111/ijfs.2025. 9358
[17] Y. Du, Intuitionistic fuzzy sets for large language model data annotation: A novel approach to side-by-side prefer ence labeling, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2505.24199
[18] Y. Fang, et al., Advancing Arabic sentiment analysis: ArSen benchmark and the improved fuzzy deep hybrid network, Proceedings of the 28th Conference on Computational Natural Language Learning, (2024), 507-516. https://doi.org/10.18653/v1/2024.conll-1.39
[19] F. Fanian, M. Kuchaki Rafsanjani, A. Borumand Saeid, An expert-aware intelligent multi-phase protocol: Metaheuristic-fuzzy-guided machine learning for enhanced generalizability in smart WRSNs, Journal of Network and Computer Applications, 236 (2026), 104464. https://doi.org/10.1016/j.jnca.2026.104464
[20] F. Fanian, M. Kuchaki Rafsanjani, M. Shokouhifar, Combined fuzzy-metaheuristic framework for bridge health monitoring using UAV-enabled rechargeable wireless sensor networks, Applied Soft Computing, 167 (2024), 112429. https://doi.org/10.1016/j.asoc.2024.112429
[21] V. Figueiredo, Fuzzy, symbolic, and contextual: Enhancing LLM instruction via cognitive scaffolding, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2508.21204
[22] J. F. Gaitán-Guerrero, et al., A novel fine-tuning and evaluation methodology for large language models on IoT raw data summaries (LLM-RawDMeth): A joint perspective in diabetes care, Computer Methods and Programs in Biomedicine, 269 (2025), 108878. https://doi.org/10.1016/j.cmpb.2025.108878
[23] R. Gauraha, A. K. Agrawal, P. Dubey, Hybrid transformer–fuzzy framework for interpretable sentiment classification in deepfake social media content, Scientific Reports, (2026). https://doi.org/10.1038/ s41598-026-58095-9
[24] C. Geng, An intelligent framework combining deep learning and fuzzy logic for accurate remote language translation, Scientific Reports, 15(1) (2025), 38736. https://doi.org/10.1038/s41598-025-22549-3
[25] S. Gilda, S. Gilda, AI-assisted engineering should track the epistemic status and temporal validity of architectural decisions, arXiv preprint, (2026). https://doi.org/10.48550/arXiv.2601.21116
[26] S. Gorle, S. B. S. Rao, P. Muthusamy, Detecting and mitigating hallucinations in large language models (LLMs) using reinforcement learning in healthcare, Journal of AI-Powered Medical Innovations, 1(1) (2024), 105-118. https://doi.org/10.60087/Japmi.Vol.03.Issue.01.Id.011
[27] W. Gu, et al., FPSO-Time: A fuzzy time series forecasting method based on fuzzy particle swarm optimization and reprogrammed large language models, IEEE Transactions on Fuzzy Systems, 34(2) (2026). https://doi.org/10. 1109/TFUZZ.2025.3637018
[28] B. Haznedar, L. Karacan, FISformer: Replacing self-attention with a fuzzy inference system in transformer models for time series forecasting, IEEE Transactions on Fuzzy Systems, (2026). https://doi.org/10.1109/TFUZZ.2026. 3690012
[29] X. Hu, et al., Fuzzy symbolic reasoning for few-shot KBQA: A CBR-inspired generative approach, International Conference on Case-Based Reasoning Research and Development, (2025), 96-110. https://doi.org/10.1007/ 978-3-031-96559-3_7
[30] Y. S. Hu, S. Marandi, M. Modarres, DML–LLM hybrid architecture for fault detection and diagnosis in sensor-rich industrial systems, Sensors, 26(6) (2026), 2008. https://doi.org/10.3390/s26062008
[31] Z. Ji, et al., Survey of hallucination in natural language generation, ACM Computing Surveys, 55(12) (2023), 1-38. https://doi.org/10.1145/3571730
[32] S. Kadavath, et al., Language models (mostly) know what they know, arXiv preprint, (2022). https://doi.org/ 10.48550/arXiv.2207.05221
[33] M. Kamal, et al., Physical fuzzy rule based unsupervised news article clustering, 2025 IEEE 49th Annual Computers, Software, and Applications Conference (COMPSAC), IEEE, (2025). https://doi.org/10.1109/COMPSAC65507. 2025.00218
[34] T. Kanagawa, Deterministic compliance failures in large language models: A structural analysis using a legacy fuzzy inference benchmark, Engineering Engrxiv Archive, (2026). https://doi.org/10.31224/6702
[35] J. Kaplan, et al., Scaling laws for neural language models, arXiv preprint, (2020). https://doi.org/10.48550/ arXiv.2001.08361
[36] J. Kim, et al., Fuzzy contrastive decoding to alleviate object hallucination in large vision-language models, Proceedings of the IEEE/CVF International Conference on Computer Vision, (2025). https://doi.org/10.1109/ ICCV51701.2025.01913
[37] M. Košprdić, et al., VerifAI: A verifiable open-source search engine for biomedical question answering, IEEE Access, 14 (2026), 45129-45147. https://doi.org/10.13140/RG.2.2.14034.82888
[38] J. Kralev, Fine-tuning LLMs for real-time fuzzy insulin control in Type I diabetes, Medicine and Pharmacology, Endocrinology and Metabolism, (2026). https://doi.org/10.20944/preprints202604.1316.v1
[39] M. Kumari, A. Chaudhary, Y. Narayan, Explainable AI (XAI): A survey of current and future opportunities, Explainable Edge AI: A Futuristic Computing Perspective, Cham: Springer International Publishing, (2022), 53-71. https://doi.org/10.1007/978-3-031-18292-1_4
[40] B. Kwon, K. Yu, Cognition-enhanced geospatial decision framework integrating fuzzy FCA, surprisingly popu lar method, and a large language model, Scientific Reports, 15(1) (2025), 23089. https://doi.org/10.1038/ s41598-025-06508-6
[41] M. Li, et al., Large language model assisted hyper-heuristic evolutionary algorithm for groundwater level prediction, Scientific Reports, 16 (2026). https://doi.org/10.1038/s41598-026-52801-3
[42] S. Lin, J. Hilton, O. Evans, TruthfulQA: Measuring how models mimic human falsehoods, in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL), 1 (2022), 3214-3252. https:// aclanthology.org/2022.acl-long.229.pdf
[43] R. Lin, Y. You, Let the fuzzy rule speak: Enhancing in-context learning debiasing with interpretability, arXiv preprint, (2024). https://doi.org/10.48550/arXiv.2412.19018
[44] L. Liu, Multi-hop reasoning and retrieval in embedding space: Leveraging large language models with knowledge, arXiv preprint, (2026). https://doi.org/10.48550/arXiv.2603.13266
[45] A. Madaan, et al., Self-refine: Iterative refinement with self-feedback, in Advances in Neural Information Processing Systems (NeurIPS), (2019), 46534-46594.
[46] A. Malinin, M. Gales, Uncertainty estimation in autoregressive structured prediction, in International Conference on Learning Representations (ICLR), (2020). https://doi.org/10.48550/arXiv.2002.07650
[47] E. H. Mamdani, S. Assilian, An experiment in linguistic synthesis with a fuzzy logic controller, International Journal of Human-Computer Studies, 51(2) (1999), 135-147. https://doi.org/10.1006/ijhc.1973.0303
[48] S. Manolache, N. Popescu, Incorporating fuzzy logic into an OpenAI-based decision-making system, UPB Scientific Bulletin, Series C: Electrical Engineering and Computer Science, 87(4) (2025).
[49] C. Molnar, Interpretable machine learning, 2nd ed., 2022. [Online]. Available: https://christophm.github.io/ interpretable-ml-book/
[50] P. Mortezaagha, A. Rahgozar, An auditable pipeline for fuzzy full-text screening in systematic reviews: Integrating contrastive semantic highlighting and LLM judgment, Artificial Intelligence Review, (2026). https://doi.org/10. 1007/s10462-026-11599-2
[51] A. Mukherjee, S. Das, ChatGPT: A fuzzy system that talks like a human, Journal of Mathematical Sciences and Computational Mathematics, 5(3) (2024). https://doi.org/10.15864/jmscm.5303
[52] T. M. Nguyen, T. T. Nguyen, T. T. Quan, A two-stage fuzzy-guided genetic algorithm for university timetabling with LLM-based preference parsing, International Conference on Multi-disciplinary Trends in Artificial Intelligence, Singapore: Springer Nature Singapore, (2025), 358-370. https://doi.org/10.1007/978-981-95-4960-3_29
[53] OpenAI, Introducing ChatGPT, OpenAI Blog, Nov. 2022. [Online]. Available: https://openai.com/blog/ chatgpt
[54] O. Orang, et al., Causal graph fuzzy LLMs: A first introduction and applications in time series forecasting, Conference: Congresso Brasileiro de Inteligência Computacional, (2025). https://doi.org/10.21528/ CBIC2025-1187477
[55] E. Özbay, F. A. Özbay, A. B. Özer, A unified AI framework for Turkish E-commerce review analysis: Sentiment classification, LLM-based summarization, and fuzzy evaluation, Applied Sciences, 16(12) (2026), 5849. https: //doi.org/10.3390/app16125849
[56] S. Qahtan, et al., Three-way decision approach based on utility and dynamic localization transformational procedures within a circular q-rung orthopair fuzzy set for ranking and grading large language models, Cognitive Computation, 17(2) (2025), 77. https://doi.org/10.1007/s12559-025-10432-2
[57] R. Radišić, S. Popov, N. Ralević, Large language model and fuzzy metric integration in assignment grading for introduction to programming type of courses, Mathematics, 14(1) (2025), 137. https://doi.org/10.3390/ math14010137
[58] J. Ren, I. Jenkinson, O. Tobora, Grounded expert systems for offshore safety: Enhancing rule-based risk assessment with retrieval-augmented generation (RAG) for auditable explanations, Conference: International Conference on Electronics Technology (ICET), (2026).
[59] M. T. Ribeiro, S. Singh, C. Guestrin, “Why should I trust you?” Explaining the predictions of any classifier, in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), (2016), 1135-1144. https://doi.org/10.1145/2939672.2939778
[60] A. Saeedi, M. Kuchaki Rafsanjani, S. Yazdani, Energy efficient clustering in IoT-based wireless sensor networks using binary whale optimization algorithm and fuzzy inference system, The Journal of Supercomputing, 81(1) (2025), 209. https://doi.org/10.1007/s11227-024-06556-1
[61] M. Sajid, et al., Fuzzy learning at 60: Future of trustworthy AI in healthcare and LLM, Conference: 2025 IEEE International Conference on Fuzzy Systems- Celebrating 60 Years of Fuzzy Sets, (2025).
[62] C. Salamea-Palacios, W. Martel-Sócala, E. Arcos-Salamea, Development of a sentiment analysis system for Chatbot responses enhanced with fuzzy logic, International Conference on Information Technology and Systems, (2025), 71 80. https://doi.org/10.1007/978-3-031-93106-2_7
[63] B. Saulnier, W. Oettgen, Fuzzy-cognitive integration and LLM-assisted interpretation: Toward traceable and cognitively faithful AI explanations, ResearchGate, (2024). https://doi.org/10.5281/zenodo.17386634
[64] R. Schuerkamp, P. J. Giabbanelli, Guiding evolutionary algorithms with large language models to learn fuzzy cognitive maps, Neural Computing and Applications, 37(18) (2025), 11891-11908. https://doi.org/10.1007/ s00521-025-11157-x
[65] K. A. Shaik, et al., A fuzzy supervisory framework for real-time optimization of robot output and LLM performance in HRI, 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI), (2025). https: //doi.org/10.1109/HRI61500.2025.10974222
[66] J. Song, et al., An uncertainty-aware framework integrating large language model and fuzzy inference system for commonsense reasoning, Expert Systems with Applications, 310 (2026), 131273. https://doi.org/10.1016/j. eswa.2026.131273
[67] T. Stefanova, S. Georgiev, Fuzzy ensemble of large language models for financial sentiment analysis, International Conference on Intelligent and Fuzzy Systems, Cham: Springer Nature Switzerland, (2025), 553-565. https://doi. org/10.1007/978-3-031-97992-7_62
[68] W. W. Tan, T. W. Chua, Uncertain rule-based fuzzy logic systems: Introduction and new directions (Mendel, JM; 2001) [book review], IEEE Computational intelligence magazine, 2(1) (2007), 72-73. https://doi.org/10.1109/ MCI.2007.357196
[69] H. Touvron, et al., LLaMA: Open and efficient foundation language models, arXiv preprint, (2023). https://doi. org/10.48550/arXiv.2302.13971
[70] K. Trinkūnas, J. Miliauskaité, Fuzzy-based assessment of stakeholder feedback in software requirements, New Trends in Computer Sciences, 3(2) (2025), 154-172. https://doi.org/10.3846/ntcs.2025.26832
[71] H. Tripathi, et al., From protocol to practice: Graded sepsis bundle compliance and actionable insights from real world ICU data, MedRxiv, (2026), 2026-04. https://doi.org/10.64898/2026.04.23.26351412
[72] A. Vaswani, et al., Attention is all you need, in Advances in Neural Information Processing Systems (NeurIPS), (2017), 5998-6008.
[73] Visual-linguistic abductive reasoning (ViLA): Integrating fuzzy scoring for multi-modal reasoning, 2026. (P063).
[74] R. Wajman, A hybrid fuzzy-sentiment framework for adaptive two-phase flow control in human-centric systems, Ad vances in Science and Technology Research Journal, 19(12) (2025), 42-55. https://doi.org/10.12913/22998624/ 209579
[75] H. Wang, et al., Collm: Industrial large-small model collaboration with fuzzy decision-making agent and self reflection, IEEE Transactions on Fuzzy Systems, 34(4) (2025). https://doi.org/10.1109/TFUZZ.2025.3594229
[76] S. Wang, et al., Fuzzy-assisted contrastive decoding improving code generation of large language models, IEEE Transactions on Fuzzy Systems, 33(8) (2025). https://doi.org/10.1109/TFUZZ.2025.3575060
[77] M. Wang, et al., SAGE: Global semantic alignment with LLMs for long-tail sequential recommendation, Proceedings of the ACM Web Conference, (2026), 6433-6444. https://doi.org/10.1145/3774904.3792456
[78] Z. Wu, Neural fuzzy logic reasoning for natural language inference, A thesis submitted in partial fulfillment of the requirements for the degree of Master of Science, (2022). https://ualberta.scholaris.ca/server/api/core/ bitstreams/748855bc-09b2-42bc-ae86-e3984a4e887c/content
[79] S. Wu, Fuzzy retrieval of power dispatching knowledge base through large language model integrated with knowledge graph, International Journal of Information Technology and Management, 24(3-4) (2025), 237-258. https://doi. org/10.1504/ijitm.2025.151550
[80] H. Wu, et al., FLAIR: Steering LLM mathematical problem solving based on a fuzzy-logic-assIsted rReasoner, Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), (2026). https://doi.org/10.18653/v1/2026.acl-long.1790
[81] Y. Xiang, et al., Large language model for secure operation of power systems, Smart Energy System Research, 1(1) (2025), 10005. https://doi.org/10.70322/sesr.2025.10005
[82] T. Xu, X. S. Tang, Electrical equipment fault diagnosis: A technique combining fuzzy logic and large language models, 2023 IEEE International Symposium on Product Compliance Engineering-Asia (ISPCE-ASIA), (2023). https://doi.org/10.1109/ISPCE-ASIA60405.2023.10365878
[83] S. Yadav, N. K. Verma, Explainable AI for sentiment classification with Type-1 fuzzy and Type-2 fuzzy models, IEEE Transactions on Fuzzy Systems, 34(7) (2026), 2239-2251. https://doi.org/10.1109/TFUZZ.2026.3688623
[84] W. Yan, C. Zhang, K. Yamada, Fuzzmonte-rag: A hybrid fuzzy logic and monte carlo approach for personalized learning optimisation in llm-assisted it education, Bulletin of Advanced Institute of Industrial Technology, 18 (2025). https://aiit.ac.jp/documents/jp/research_collab/research/bulletin/18th/03_Yan.pdf
[85] H. Yang, et al., Generative fuzzy system for sequence generation, arXiv preprint, (2024). https://doi.org/10. 48550/arXiv.2411.1386
[86] T. Yao, et al., Multi-agent fuzzy reinforcement learning with LLM for cooperative navigation of endovascular robotics, IEEE Transactions on Fuzzy Systems, 34(4) (2026), 1109-1119. https://doi.org/10.1109/TFUZZ.2025. 3585934
[87] T. Yoshida, Y. Sueoka, K. Osuka, Verification of a two-step inference model for cooperative evaluation of robot actions using foundation models, International Symposium on Distributed Autonomous Robotic Systems, Cham: Springer Nature Switzerland, (2024), 395-410. https://doi.org/10.1007/978-3-032-04584-3_27
[88] L. Yousofvand, M. B. Dowlatshahi, M. Pirdadeh Beiranvand, Pneumonia detection in chest X-ray images using Convolutional Neural Network and fuzzy VIKOR, Iranian Journal of Fuzzy Systems, 22(5) (2025), 159-179. https: //doi.org/10.22111/ijfs.2025.51603.9119
[89] L. A. Zadeh, Fuzzy sets, Information and Control, 8(3) (1965), 338-353. https://doi.org/10.1016/ S0019-9958(65)90241-X
[90] L. A. Zadeh, The concept of a linguistic variable and its application to approximate reasoning, Information Sciences, 8(3) (1975), 199-249.https://doi.org/10.1016/0020-0255(75)90036-5
[91] B. Zhang, et al., Large language model enhanced fuzzy logic fusion framework for stance detection, CSIG Conference on Emotional Intelligence, Singapore: Springer Nature Singapore, (2024), 130-144. https://doi.org/10.1007/ 978-981-96-5084-2_9
[92] Y. Zhang, et al., Siren’s song in the AI ocean: A survey on hallucination in large language models, Computational Linguistics, 51(4), (2025). https://doi.org/10.1162/coli.a.16
[93] M. Zhang, et al., Bridging stochasticity and fuzziness: Automated construction of triangular fuzzy numbers via LLM temperature sampling for managerial decision support, Information, 17(4) (2026), 349. https://doi.org/ 10.3390/info17040349
[94] Z. Zhang, C. Valeo, Fuzzy-based input method for uncertainty quantification in a deterministic model comparison with ChatGPT for peak flow prediction, Journal of Hydrology X, 28-29 (2025), 100208. https://doi.org/10. 1016/j.hydroa.2025.100208
[95] W. Zheng, et al., Llm-as-a-fuzzy-judge: Fine-tuning large language models as a clinical evaluation judge with fuzzy logic, arXiv preprint, (2025). https://doi.org/10.48550/arXiv.2506.11221