Shufan Ming's Dissertation Defense
PhD candidate Shufan Ming will present her dissertation titled, "Enhancing Accessibility of Biomedical Literature with Knowledge-Guided Large Language Models." Her final examination committee includes Associate Professor Halil Kilicoglu (chair), Assistant Professor Yue Guo, Professor Bertram Ludäscher, and Associate Professor Vetle Torvik.
Abstract
Access to biomedical knowledge is crucial for transforming information in the literature into accessible and actionable insights that improve clinical applications, biomedical research, and public health literacy. With the rapid growth in biomedical publications, both researchers and the general public face challenges in processing the vast amount of biomedical literature, leading to information overload. Researchers and clinicians need efficient methods to extract and synthesize key information for evidence-based decision-making, while non-expert audiences require understandable, layperson-friendly biomedical knowledge to improve health literacy and support informed decision-making.
Natural language processing (NLP) techniques, particularly large language models (LLMs), have emerged as promising tools for biomedical information extraction and synthesis. However, LLMs face several limitations: they often lack explicit semantic grounding, struggle with domain-specific reasoning, and exhibit issues such as hallucinations and limited interpretability. These models primarily rely on statistical co-occurrence patterns in large text corpora, which can lead to a shallow representation of domain-specific biomedical knowledge.
Biomedical domain knowledge is represented in highly structured and semantically rich resources, such as the Unified Medical Language System (UMLS), Medical Subject Headings (MeSH), and other knowledge sources that explicitly define biomedical concepts and their relationships. These symbolic resources can complement the statistical learning mechanisms of neural language models, creating a synergistic framework that can enhance the accuracy, interpretability, and usability of biomedical NLP applications.
This dissertation explores how integrating structured biomedical knowledge—such as ontologies and controlled vocabularies—into transformer-based language models can enhance predictive performance, interpretability, and reliability in biomedical NLP applications. Specifically, it focuses on three key tasks: extracting relationships between entities from biomedical literature to support biomedical research applications, such as literature-based discovery; summarizing biomedical abstracts into layperson-friendly summaries to improve health literacy; and classifying biomedical publication types to improve automatic literature indexing and support evidence synthesis for clinical decision-making.
Across these tasks, the findings demonstrate that structured biomedical knowledge can complement neural language models in distinct roles across training and inference. By combining the contextual modeling capabilities of neural language models with explicit biomedical semantics and constraints, this dissertation provides insights into the development of more grounded, interpretable, and reliable systems for extracting, synthesizing, and communicating biomedical knowledge.
Questions? Contact Shufan Ming.