Large Language Models in Biomedical Knowledge Extraction

Authors

  • Andreas Horvath Author
  • Anna Ivanov Author
  • Andreas Bianchi Author

DOI:

https://doi.org/10.5281/zenodo.19549529

Keywords:

: Large language models; Biomedical knowledge extraction; Named entity recognition; Relation extraction; Retrieval-augmented generation; Knowledge graphs; Drug-drug interactions; Hallucination reduction; PubMed mining

Abstract

Biomedical literature doubles every 3-4 years, with over 1.5 million PubMed-indexed articles published annually, creating an insurmountable knowledge extraction bottleneck for researchers, clinicians, and drug developers. Large language models (LLMs) trained on biomedical corpora can automate extraction of structured knowledge -- drug-target interactions, gene-disease associations, adverse event reports, and clinical trial outcomes -- from unstructured text at scale. However, LLM-based extraction suffers from hallucination (generating plausible but factually incorrect relations), incomplete grounding (assertions without traceable evidence), and inconsistent performance across biomedical subdomains. This study developed and benchmarked six knowledge extraction approaches -- rule-based pattern matching (SemRep), BioBERT named entity recognition + relation extraction, PubMedBERT fine-tuned pipeline, GPT-4 few-shot prompting, domain-adapted BioMistral, and a proposed retrieval-augmented biomedical extraction system (RABES) combining a fine-tuned BioMistral-7B backbone with dense passage retrieval from PubMed, structured knowledge graph grounding (UMLS + DrugBank + STRING), and chain-of-evidence generation with citation tracking - across five biomedical relation extraction tasks: drug-drug interactions (DDI, n = 4,892 pairs), gene-disease associations (GDA, n = 12,486), adverse drug events (ADE, n = 8,248), protein-protein interactions (PPI, n = 6,842), and chemical-protein interactions (CPI, n = 5,624). RABES achieved the highest micro-F1 across all tasks: DDI 0.892 (vs GPT-4 0.824, PubMedBERT 0.846; p < 0.001), GDA 0.868, ADE 0.878, PPI 0.842, CPI 0.856 -- mean F1 0.867 versus 0.812 for GPT-4 and 0.798 for PubMedBERT. Critically, RABES reduced hallucination rate from 14.8% (GPT-4 few-shot) to 2.4% through knowledge graph grounding, and provided traceable evidence chains (cited PubMed abstracts) for 96.2% of extracted relations. These results establish retrieval-augmented, knowledge-grounded LLMs as the state-of-the-art for biomedical knowledge extraction.

Downloads

Published

2026-08-15

How to Cite

Large Language Models in Biomedical Knowledge Extraction. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 3(4), 110-118. https://doi.org/10.5281/zenodo.19549529

Similar Articles

1-10 of 82

You may also start an advanced similarity search for this article.