Large Language Models in Biomedical Knowledge Extraction
DOI:
https://doi.org/10.5281/zenodo.19549529Keywords:
: Large language models; Biomedical knowledge extraction; Named entity recognition; Relation extraction; Retrieval-augmented generation; Knowledge graphs; Drug-drug interactions; Hallucination reduction; PubMed miningAbstract
Biomedical literature doubles every 3-4 years, with over 1.5 million PubMed-indexed articles published annually, creating an insurmountable knowledge extraction bottleneck for researchers, clinicians, and drug developers. Large language models (LLMs) trained on biomedical corpora can automate extraction of structured knowledge -- drug-target interactions, gene-disease associations, adverse event reports, and clinical trial outcomes -- from unstructured text at scale. However, LLM-based extraction suffers from hallucination (generating plausible but factually incorrect relations), incomplete grounding (assertions without traceable evidence), and inconsistent performance across biomedical subdomains. This study developed and benchmarked six knowledge extraction approaches -- rule-based pattern matching (SemRep), BioBERT named entity recognition + relation extraction, PubMedBERT fine-tuned pipeline, GPT-4 few-shot prompting, domain-adapted BioMistral, and a proposed retrieval-augmented biomedical extraction system (RABES) combining a fine-tuned BioMistral-7B backbone with dense passage retrieval from PubMed, structured knowledge graph grounding (UMLS + DrugBank + STRING), and chain-of-evidence generation with citation tracking - across five biomedical relation extraction tasks: drug-drug interactions (DDI, n = 4,892 pairs), gene-disease associations (GDA, n = 12,486), adverse drug events (ADE, n = 8,248), protein-protein interactions (PPI, n = 6,842), and chemical-protein interactions (CPI, n = 5,624). RABES achieved the highest micro-F1 across all tasks: DDI 0.892 (vs GPT-4 0.824, PubMedBERT 0.846; p < 0.001), GDA 0.868, ADE 0.878, PPI 0.842, CPI 0.856 -- mean F1 0.867 versus 0.812 for GPT-4 and 0.798 for PubMedBERT. Critically, RABES reduced hallucination rate from 14.8% (GPT-4 few-shot) to 2.4% through knowledge graph grounding, and provided traceable evidence chains (cited PubMed abstracts) for 96.2% of extracted relations. These results establish retrieval-augmented, knowledge-grounded LLMs as the state-of-the-art for biomedical knowledge extraction.Downloads
Published
2026-08-15
Issue
Section
Articles
How to Cite
Large Language Models in Biomedical Knowledge Extraction. (2026). The Biosis Bulletin: Bioscience and Information Science Journal , 3(4), 110-118. https://doi.org/10.5281/zenodo.19549529

