A Multi-Agent Large Language Model Framework for Interpretable Multi-Omics Literature Mining and Gene–Disease Hypothesis Generation

Authors

  • Emily(Bingbing) Chen Department of Biomedical Informatics, University of Pittsburgh, Pittsburgh, PA, USA

Keywords:

large language models; multi-agent systems; multi-omics; literature mining; gene– disease association; biomedical hypothesis generation; interpretability

Abstract

The biomedical literature contains a rapidly expanding body of evidence linking genomic vari- ation, transcriptomic regulation, proteomic interaction, epigenetic modification, metabolomic state, and disease phenotypes. However, conventional literature-mining systems typically opti- mize isolated extraction tasks and provide limited support for interpretable, cross-omics hypoth- esis generation. This paper presents OmicsAgent, a multi-agent large language model (LLM) framework for interpretable multi-omics literature mining and gene–disease hypothesis gener- ation. The framework decomposes biomedical discovery into coordinated subtasks performed by specialized agents for retrieval, entity normalization, omics evidence extraction, relation ver- ification, mechanistic synthesis, and hypothesis ranking. Inspired by structured multi-agent collaboration for complex task solving, particularly the role-based and verifier-guided design discussed by Wang, S., Feng, Y., & Fang, X. (2026). A Large Language Model-Enabled Multi- Agent Collaboration Method for Complex Task Solving., OmicsAgent replaces unrestricted agent dialogue with typed semantic messages, evidence-grounded memory, and reliability-weighted de- cision fusion. We evaluate the framework on a curated corpus of 128,640 PubMed abstracts, 18,420 PubMed Central full-text sections, and benchmark annotations derived from DisGeNET, OMIM, GWAS Catalog, CTD, UniProt, and Reactome. Compared with a single-LLM retrieval- augmented baseline, OmicsAgent improves gene–disease relation F1 from 0.743 to 0.842, in- creases novel hypothesis precision at 50 from 0.312 to 0.428, and reduces unsupported mecha- nistic claims by 34.6%. Ablation studies show that verifier agents, omics-specific specialization, and evidence-constrained fusion are all necessary for robust performance. The results indicate that multi-agent LLM systems can support transparent biomedical knowledge synthesis when their collaboration protocol is explicitly constrained by curated identifiers, provenance traces, and verifiable intermediate states.

References

1. Wang, S., Feng, Y., & Fang, X. (2026, May). A Large Language Model-Enabled Multi-Agent Collaboration Method for Complex Task Solving. In 2026 6th International Symposium on Computer Technology and Information Science (ISCTIS) (pp. 253-256). IEEE.

2. Piero, J., et al. (2020). The DisGeNET knowledge platform for disease genomics. Nucleic Acids Research, 48(D1), D845–D855.

3. Amberger, J. S., Bocchini, C. A., Scott, A. F., & Hamosh, A. (2019). OMIM.org: leveraging knowledge across phenotype–gene relationships. Nucleic Acids Research, 47(D1), D1038–D1043.

4. Sollis, E., et al. (2023). The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Research, 51(D1), D977–D985.

5. Davis, A. P., et al. (2023). Comparative Toxicogenomics Database: update 2023. Nucleic Acids Research, 51(D1), D1257–D1262.

6. Szklarczyk, D., et al. (2023). The STRING database in 2023: protein–protein association networks and functional enrichment analyses. Nucleic Acids Research, 51(D1), D638–D646.

7. Gu, Y., et al. (2021). Domain-specific language model pretraining for biomedical natural lan- guage processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23.

8. Lee, J., et al. (2020). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240.

9. Brown, T. B., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

10. Wei, J., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837.

11. Lewis, P., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

12. Yao, S., et al. (2023). ReAct: Synergizing reasoning and acting in language models. Interna- tional Conference on Learning Representations.

13. Shinn, N., et al. (2023). Reflexion: Language agents with verbal reinforcement learning. Ad- vances in Neural Information Processing Systems, 36.

14. Du, Y., et al. (2023). Improving factuality and reasoning in language models through multiagent debate. arXiv preprint arXiv:2305.14325.

15. Wu, Q., et al. (2024). AutoGen: Enabling next-gen LLM applications via multi-agent conver- sation. Conference on Language Modeling.

16. Hong, S., et al. (2024). MetaGPT: Meta programming for a multi-agent collaborative frame- work. International Conference on Learning Representations.

17. Li, G., et al. (2023). CAMEL: Communicative agents for mind exploration of large language model society. Advances in Neural Information Processing Systems, 36.

18. Chen, Q., Lee, K., Yan, S., Kim, S., Wei, C. H., & Lu, Z. (2020). BioConceptVec: Creating and evaluating literature-based biomedical concept embeddings on a large scale. PLOS Computa- tional Biology, 16(4), e1007617.

19. Beltagy, I., Lo, K., & Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. Proceedings of EMNLP-IJCNLP, 3615–3620.

20. Luo, R., et al. (2022). BioGPT: Generative pre-trained transformer for biomedical text gener- ation and mining. Briefings in Bioinformatics, 23(6), bbac409.

21. Singhal, K., et al. (2023). Large language models encode clinical knowledge. Nature, 620, 172–180.

22. Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., & Lu, X. (2019). PubMedQA: A dataset for biomedical research question answering. Proceedings of EMNLP-IJCNLP, 2567–2577.

23. Bodenreider, O. (2004). The Unified Medical Language System: Integrating biomedical termi- nology. Nucleic Acids Research, 32(Database issue), D267–D270.

24. Ashburner, M., et al. (2000). Gene Ontology: Tool for the unification of biology. Nature Genetics, 25, 25–29.

25. Gillespie, M., et al. (2022). The Reactome pathway knowledgebase 2022. Nucleic Acids Re- search, 50(D1), D687–D692.

26. Wishart, D. S., et al. (2022). HMDB 5.0: The Human Metabolome Database for 2022. Nucleic Acids Research, 50(D1), D622–D631.

27. Ko…hler, S., et al. (2021). The Human Phenotype Ontology in 2021. Nucleic Acids Research, 49(D1), D1207–D1217.

28. Oughtred, R., et al. (2021). The BioGRID database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Science, 30(1), 187–2

Downloads

Published

2026-08-27

How to Cite

Emily(Bingbing) Chen. (2026). A Multi-Agent Large Language Model Framework for Interpretable Multi-Omics Literature Mining and Gene–Disease Hypothesis Generation. Bioinformatics Insights and Analytics, 1(1). Retrieved from https://www.bioinfia.org/index.php/home/article/view/188