CONFIDENCE-DRIVEN GRAPH-OF-THOUGHT RETRIEVAL-AUGMENTED AGENTS FOR VERIFIABLE VULNERABILITY PATCH GENERATION IN MULTI-LANGUAGE LEGACY CODEBASES

Authors

  • Nibras Hmaizah General Directorate of Education in Al- Qadisiyah Govemorate, Diwaniyah, Iraq
  • Roaa Ghanim General Directorate of Education in Al- Qadisiyah Govemorate, Diwaniyah, Iraq

Keywords:

Automated vulnerability repair, large language models, retrieval-augmented generation, graph-of-thought reasoning, uncertainty calibration, selective prediction, legacy software, software security

Abstract

It is much more difficult to automatically repair security vulnerabilities in legacy software that is multi-language, than it is to detect them, for an automated patch that successfully compiles does not necessarily improve the security, functionality, or lack of intrusiveness of the software. Most current large language-model (LLM) repair systems are prone to producing grammatically correct but unverified patches, lacking evidence of what edits it has made, and giving inaccurate confidence estimates, and are rarely tested across languages under leakage control. We introduce UC-GoT-RAG, a framework consisting of an agentic architecture, vulnerability-aware hybrid retrieval, an evidence-linked Graph-of-Thought (GoT) reasoning representation, specialised analysis/generation/verification agents, and a patch-specific uncertainty-decomposition and calibration mechanism governing selective acceptance, abstention, or human escalation. Each patch is put through a rigorous staged verification process including application, build, functional and regression testing, neutralising the exploit or proof of vulnerability, security testing, and static analysis, ensuring correctness is executable, not textual. On a leakage-controlled five-language corpus collected from public commits which fixed CVE bugs, UC-GoT-RAG achieves a Verifiable Correct Patch Rate of 46.9%, which is 9.7 points higher than the best controlled baseline, and also drops the Expected Calibration Error from 0.164 to 0.041, reducing incorrect-patch acceptance at a fixed coverage by half. The most significant gains in retrieval and graph reasoning are in rare CWE families and high-legacy-burden projects. The benchmark split, prompts and configurations are released along with the patch validation harness for reproduction.

References

M. Fu, C. Tantithamthavorn, T. Le, V. Nguyen, and D. Phung, "VulRepair: A T5-based automated software vulnerability repair," in Proc. 30th ACM Joint Eur. Softw. Eng. Conf. Symp. Found. Softw. Eng. (ESEC/FSE), 2022, pp. 935–947.

Z. Chen, S. Kommrusch, and M. Monperrus, "Neural transfer learning for repairing security vulnerabilities in C code," IEEE Trans. Softw. Eng., vol. 49, no. 1, pp. 147–165, 2023.

X. Zhou, K. Kim, B. Xu, D. Han, and D. Lo, "Out of sight, out of mind: Better automatic vulnerability repair by broadening input ranges and sources," in Proc. IEEE/ACM 46th Int. Conf. Softw. Eng. (ICSE), 2024, pp. 1071–1083.

G. Bhandari, A. Naseer, and L. Moonen, "CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software," in Proc. 17th Int. Conf. Predictive Models Data Anal. Softw. Eng. (PROMISE), 2021, pp. 30–39.

Z. Wei et al., "PATCHEVAL: A new benchmark for evaluating LLMs on patching real-world vulnerabilities," arXiv:2511.11019, 2025.

M. Besta et al., "Graph of Thoughts: Solving elaborate problems with large language models," in Proc. AAAI Conf. Artif. Intell., vol. 38, no. 16, 2024, pp. 17682–17690.

S. Yao et al., "Tree of Thoughts: Deliberate problem solving with large language models," in Adv. Neural Inf. Process. Syst. (NeurIPS), 2023.

J. Wei et al., "Chain-of-thought prompting elicits reasoning in large language models," in Adv. Neural Inf. Process. Syst. (NeurIPS), 2022.

X. Wang et al., "Self-consistency improves chain of thought reasoning in language models," in Proc. Int. Conf. Learn. Represent. (ICLR), 2023.

[10] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Adv. Neural Inf. Process. Syst. (NeurIPS), 2020.

C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On calibration of modern neural networks," in Proc. 34th Int. Conf. Mach. Learn. (ICML), 2017, pp. 1321–1330.

M. Minderer et al., "Revisiting the calibration of modern neural networks," in Adv. Neural Inf. Process. Syst. (NeurIPS), 2021.

A. N. Angelopoulos and S. Bates, "A gentle introduction to conformal prediction and distribution-free uncertainty quantification," arXiv:2107.07511, 2021.

S. Kadavath et al., "Language models (mostly) know what they know," arXiv:2207.05221, 2022.

L. Kuhn, Y. Gal, and S. Farquhar, "Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation," in Proc. Int. Conf. Learn. Represent. (ICLR), 2023.

Y. Geifman and R. El-Yaniv, "Selective classification for deep neural networks," in Adv. Neural Inf. Process. Syst. (NeurIPS), 2017.

M. Chen et al., "Evaluating large language models trained on code," arXiv:2107.03374, 2021.

C. E. Jimenez et al., "SWE-bench: Can language models resolve real-world GitHub issues?" in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.

J. Yang et al., "SWE-agent: Agent-computer interfaces enable automated software engineering," in Adv. Neural Inf. Process. Syst. (NeurIPS), 2024.

Y. Zhang, H. Ruan, Z. Fan, and A. Roychoudhury, "AutoCodeRover: Autonomous program improvement," in Proc. 33rd ACM SIGSOFT Int. Symp. Softw. Testing Anal. (ISSTA), 2024.

C. S. Xia, Y. Deng, S. Dunn, and L. Zhang, "Agentless: Demystifying LLM-based software engineering agents," arXiv:2407.01489, 2024.

I. Bouzenia, P. Devanbu, and M. Pradel, "RepairAgent: An autonomous, LLM-based agent for program repair," in Proc. IEEE/ACM 47th Int. Conf. Softw. Eng. (ICSE), 2025.

U. Kulsum, H. Zhu, B. Xu, and M. d’Amorim, "A case study of LLM for automated vulnerability repair: Assessing impact of reasoning and patch validation feedback," in Proc. 1st ACM Int. Conf. AI-Powered Softw. (AIware), 2024, pp. 103–111.

H. Pearce, B. Tan, B. Ahmad, R. Karri, and B. Dolan-Gavitt, "Examining zero-shot vulnerability repair with large language models," arXiv:2112.02125, 2021.

Y. Nong, H. Yang, L. Cheng, H. Hu, and H. Cai, "APPATCH: Automated adaptive prompting large language models for real-world software vulnerability patching," arXiv:2408.13597, 2024.

Y. Li, F. H. Shezan, B. Wei, G. Wang, and Y. Tian, "SoK: Towards effective automated vulnerability repair," arXiv:2501.18820, 2025.

J. Fan, Y. Li, S. Wang, and T. N. Nguyen, "A C/C++ code vulnerability dataset with code changes and CVE summaries," in Proc. 17th Int. Conf. Mining Softw. Repositories (MSR), 2020, pp. 508–512.

N. Jiang, T. Lutellier, and L. Tan, "CURE: Code-aware neural machine translation for automatic program repair," in Proc. IEEE/ACM 43rd Int. Conf. Softw. Eng. (ICSE), 2021, pp. 1161–1173.

N. Jiang, K. Liu, T. Lutellier, and L. Tan, "Impact of code language models on automated program repair," in Proc. IEEE/ACM 45th Int. Conf. Softw. Eng. (ICSE), 2023, pp. 1430–1442.

C. Le Goues, T. Nguyen, S. Forrest, and W. Weimer, "GenProg: A generic method for automatic software repair," IEEE Trans. Softw. Eng., vol. 38, no. 1, pp. 54–72, 2012.

M. P. Naeini, G. F. Cooper, and M. Hauskrecht, "Obtaining well calibrated probabilities using Bayesian binning," in Proc. AAAI Conf. Artif. Intell., 2015, pp. 2901–2907.

N. Nashid, M. Sintaha, and A. Mesbah, "Retrieval-based prompt selection for code-related few-shot learning," in Proc. IEEE/ACM 45th Int. Conf. Softw. Eng. (ICSE), 2023, pp. 2450–2462.

P. Manakul, A. Liusie, and M. J. F. Gales, "SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models," in Proc. Conf. Empir. Methods Nat. Lang. Process. (EMNLP), 2023.

Q. Zhang et al., "A survey of learning-based automated program repair," ACM Trans. Softw. Eng. Methodol., vol. 33, no. 2, pp. 1–69, 2024.

H. Lee, Z. Zhang, H. Lu, and L. Zhang, "SEC-bench: Automated benchmarking of LLM agents on real-world software security tasks," arXiv:2506.11791, 2025.

Downloads

Published

2026-09-04

How to Cite

Hmaizah, N., & Ghanim, R. (2026). CONFIDENCE-DRIVEN GRAPH-OF-THOUGHT RETRIEVAL-AUGMENTED AGENTS FOR VERIFIABLE VULNERABILITY PATCH GENERATION IN MULTI-LANGUAGE LEGACY CODEBASES. Kyzylorda Scholarly Review, 3(4), 1–21. Retrieved from https://bulletin.ouk.kz/index.php/bulletin/article/view/95