This is an unedited manuscript accepted for publication and provided as an Article in Press for early access at the author’s request. The article will undergo copyediting, typesetting, and galley proof review before final publication. Please be aware that errors may be identified during production that could affect the content. All legal disclaimers of the journal apply.
Rishabh Kumar,
Anandu S,
Aman Kumar Jha,
- Student, Greater Noida Institute of Technology (IPU), Uttar Pradesh, India
- Student, GNIT (IPU), Knowledge Park 2, Greater Noida, 201310, Uttar Pradesh, India
- Student, GNIT (IPU), Knowledge Park 2, Greater Noida, 201310, Uttar Pradesh, India
Abstract
The exponential growth of multi-omics data and the increasing complexity of disease-associated protein interactomes have rendered conventional drug target identification pipelines computationally and epistemologically inadequate. This paper presents the Neuro-Symbolic Agentic AI for Scientific Discovery (NS-AASD) framework, a unified architecture that cohesively integrates deep reinforcement learning (DRL) exploration strategies, variational quantum simulation (VQS) of protein conformational dynamics, and XAI-audited large language model (LLM) hypothesis generation within an autonomous scientific discovery loop. The NS-AASD agent operates over a heterogeneous biomedical knowledge graph comprising 4.7 million nodes and 23.1 million typed edges, employing a proximal policy optimization (PPO) variant augmented with graph attention encoders and curiosity-driven intrinsic rewards to navigate the hypothesis space. Quantum simulation subroutines—implemented via parameterized quantum circuits on 52-qubit hardware and extended through tensor-network emulation—resolve protein binding-pocket dynamics at quantum mechanical accuracy, supplying physics-grounded docking affinity estimates that constrain LLM hypothesis plausibility. SHAP attribution auditing, counterfactual explanation generation, and ontology-aligned concept probing constitute the XAI audit stratum, rendering all generated hypotheses traceable to supporting evidence and biologically interpretable. Evaluated against four benchmark oncology target identification tasks—KRAS G12C allosteric pocket mapping, CDK4/6 selectivity differentiation, TEAD transcriptional coactivator druggability assessment, and IDH1/IDH2 isoform- selective inhibitor design—NS-AASD achieves a mean Hypothesis Quality Score (HQS) of 0.912 ± 0.009, surpassing the best existing neuro-symbolic baseline by 38.4% and reducing hypothesis-to- validated-lead cycle time by an estimated 73.1%. These results substantiate NS-AASD as a transformative paradigm for AI-augmented pharmaceutical discovery.
Keywords: neuro-symbolic AI, drug target identification, deep reinforcement learning, quantum simulation, explainable AI, LLM hypothesis generation, knowledge graph, protein conformational dynamics, SHAP attribution, autonomous scientific discovery.
References
[1] Paul SM, Mytelka DS, Dunwiddie CT, Persinger CC, Munos BH, Lindborg SR, Schacht AL. How to improve R&D productivity: the pharmaceutical industry’s grand challenge. Nature reviews Drug discovery. 2010 Mar;9(3):203-14.
[2] Hamilton W, Ying Z, Leskovec J. Inductive representation learning on large graphs. Advances in neural information processing systems. 2017;30.
[3] Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. 2017 Jul 20.
[4] Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Advances in neural information processing systems. 2017;30.
[5] Peruzzo A, McClean J, Shadbolt P, Yung MH, Zhou XQ, Love PJ, Aspuru-Guzik A, O’brien JL. A variational eigenvalue solver on a photonic quantum processor. Nature communications. 2014 Jul 23;5(1):4213.
[6] Hasin Y, Seldin M, Lusis A. Multi-omics approaches to disease. Genome biology. 2017 May 5;18(1):83.
[7] Zitnik M, Agrawal M, Leskovec J. Modeling polypharmacy side effects with graph convolutional networks. Bioinformatics. 2018 Jul 1;34(13):i457-66.
[8] Pathak D, Agrawal P, Efros AA, Darrell T. Curiosity-driven exploration by self-supervised prediction. InInternational conference on machine learning 2017 Jul 17 (pp. 2778-2787). PMLR.
[9] Cao, D.S., Xu, Q.S., & Liang, Y.Z. (2022). Deep reinforcement learning in drug discovery. Expert Opinion on Drug Discovery, 17(5), 521–535.
[10] Kim B, Wattenberg M, Gilmer J, Cai C, Wexler J, Viegas F. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). InInternational conference on machine learning 2018 Jul 3 (pp. 2668- 2677). PMLR.
[11] Xu M, Wang H, Ni B, Guo H, Tang J. Self-supervised graph-level representation learning with local and global structure. InInternational conference on machine learning 2021 Jul 1 (pp. 11548-11558). PMLR.
[12] Jumper J, Evans R, Pritzel A, Green T, Figurnov M, Ronneberger O, Tunyasuvunakool K, Bates R, Žídek A, Potapenko A, Bridgland A. Highly accurate protein structure prediction with AlphaFold. nature. 2021 Aug 26;596(7873):583-9.
[13] Rubin NC, Berry DW, Malone FD, White AF, Khattar T, DePrince III AE, Sicolo S, Küehn M, Kaicher M, Lee J, Babbush R. Fault-tolerant quantum simulation of materials using Bloch orbitals. PRX Quantum. 2023 Oct 1;4(4):040303.
[14] Kettle JG, Bagal SK, Bickerton S, Bodnarchuk MS, Breed J, Carbajo RJ, et al. Structure-Based Design and Pharmacokinetic Optimization of Covalent Allosteric Inhibitors of the Mutant GTPase KRASG12C. Journal of Medicinal Chemistry. 2020 Feb 5;63(9):4468–83.
[15] Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W. Lora: Low-rank adaptation of large language models. Iclr. 2022 Apr 25;1(2):3.
[16] Wachter S, Mittelstadt B, Russell C. Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harv. JL & Tech.. 2017;31:841.
[17] Iten R, Metger T, Wilming H, Del Rio L, Renner R. Discovering physical concepts with neural networks. Physical review letters. 2020 Jan 10;124(1):010508.
[18] Touvron H, Martin L, Stone K, Albert P, Almahairi A, Babaei Y, Bashlykov N, Batra S, Bhargava P, Bhosale S, Bikel D. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. 2023 Jul 18.
[19] Jörg Rahnenführer, De Bin R, Benner A, Ambrogi F, Lusa L, Boulesteix AL, et al. Statistical analysis of high- dimensional biomedical data: a gentle introduction to analytical goals, common approaches and challenges. BMC Medicine. 2023 May 15;21(1):182–2. Available from: https://pmc.ncbi.nlm.nih.gov/articles/PMC10186672/
[20] EU Commission. Proposal for a regulation laying down harmonised rules on artificial intelligence. Brussels. 2021;21:2021.
[21] Smaldone AM, Shee Y, Kyro GW, Xu C, Vu NP, Dutta R, et al. Quantum Machine Learning in Drug Discovery: Applications in Academia and Pharmaceutical Industries. arXiv (Cornell University). 2024 Sept 24;. Available from: https://www.researchgate.net/publication/384295321_Quantum_Machine_Learning_in_Drug_Discovery_Applications_i n_Academia_and_Pharmaceutical_Industries
[22] Dara S, Dhamercherla S, Jadav SS, Babu CM, Ahsan MJ. Machine learning in drug discovery: a review. Artificial intelligence review. 2022 Mar;55(3):1947-99.

Research and Reviews : Journal of Computational Biology
| Volume | 15 | |
| 02 | ||
| Received | 13/07/2026 | |
| Accepted | 16/07/2026 | |
| Published | 26/07/2026 | |
| Publication Time | 13 Days |