Graduation Year

2026

Document Type

Dissertation

Degree

Ph.D.

Degree Name

Doctor of Philosophy (Ph.D.)

Degree Granting Department

Computer Science and Engineering

Major Professor

John Licato, Ph.D.

Committee Member

Ankur Mali, Ph.D.

Committee Member

Gene Louis Kim, Ph.D.

Committee Member

Trung Le, Ph.D.

Committee Member

Thanh Thieu, Ph.D.

Keywords

NLP, LLM in Oncology, Tumor Phenotype and Disease Progression Modeling

Abstract

The rapid growth of electronic health records (EHRs) has created new opportunities to apply machine learning to clinical data.However, a large portion of important clinical information is still stored in unstructured text, such as pathology reports, radiology reports, and longitudinal clinical notes.These documents contain key details about tumor characteristics, diagnoses, treatments, and patient outcomes.Extracting structured and useful information from this text is challenging due to complex medical language, varied document formats, and the need to combine information across multiple reports over time.This dissertation studies how language models can be designed and adapted to better extract and use oncology-specific information from clinical text. The work is organized into three tasks that increase in complexity. First, I study how to improve question answering by integrating structured clinical knowledge into generative language models. I develop a pipeline that uses encoder-based models to extract important clinical information and include it in prompts. This approach helps the generative model produce more accurate and clinically meaningful answers. Next, I extend my work to tumor phenotype extraction from pathology reports. This task is difficult because pathology reports are long and contain detailed clinical information. To address this, I evaluate different transformer architectures and study how well they handle long clinical text. Finally, I focus on longitudinal clinical reasoning by studying disease progression detection from patient timelines. This task is more complex because it requires combining information from multiple reports over time. I evaluate how well language models can identify progression events and study their performance under real-world class imbalance.

Overall, this dissertation shows that language models can be improved by injecting structured clinical information into the generation process, leading to more accurate and clinically meaningful outputs.To support this, the second and third parts focus on identifying the most effective ways to extract high-quality clinical information from complex pathology reports and longitudinal patient data. The tumor phenotype extraction task determines the best model architectures for capturing detailed clinical attributes, while the progression detection task extends this by identifying key clinical events over time. These components are closely connected, as better information extraction directly strengthens the knowledge that is fed back into the generative model. Together, this work demonstrates that aligning information extraction and generation enables smaller, specialized models to achieve strong performance while remaining practical for real-world, privacy-sensitive oncology applications.

Share

COinS