Publication Date

2025

Document Type

Dissertation

Committee Members

Michael Raymer, Ph.D. (Advisor); Katherine Winner, M.D. (Committee Member); Cogan Shimizu, Ph.D. (Committee Member); Tanvi Banerjee, Ph.D. (Committee Member)

Degree Name

Doctor of Philosophy (PhD)

Abstract

The widespread adoption of electronic medical records has created a vast reservoir of clinical data that can be leveraged to better understand how interventions relate to patient outcomes. Much of this information, however, exists as unstructured free-text, posing significant challenges for traditional statistical and machine-learning methods. Solving these challenges would allow the extraction of specific patient subpopulations (clinically relevant cohorts of individuals who share overlapping symptoms, risk factors, or diagnostic criteria), which could be used in precision medicine. Despite this promise, extracting these subpopulations from unstructured medical notes is an ongoing challenge due to the variability of clinical language and the complex nature of patient conditions. We here describe and demonstrate a pipeline that combines named entity recognition, transformer embeddings, guided dimensionality reduction, and retrieval-augmented generation knowledge graph integration to enhance patient extraction. This pipeline unites multiple approaches to natural language processing: extraction with an ontologically-supported named entity recognition tool, term embedding with a biomedically-trained transformer, encoder-mediated dimensionality paring, and retrieval-augmented generation-assisted knowledge graph construction. Finally, these elements are united with an attention-gated feed forward neural network. We show that this design is accurate, , and configurability. We evaluate the pipeline on multiple clinically-relevant datasets, with experiments demonstrating improvements over LLMs, including GPT 4o, in classifying medical reports by specialty and mental health relevance. Our results show that incorporating knowledge graphs and dimensionality reduction enhances precision and interpretability while maintaining adaptability for different research queries.

Page Count

162

Department or Program

Department of Computer Science and Engineering

Year Degree Awarded

2025


Share

COinS