Graduation Year

2026

Document Type

Dissertation

Degree

Ph.D.

Degree Name

Doctor of Philosophy (Ph.D.)

Degree Granting Department

Electrical Engineering

Major Professor

Ghulam Rasool, Ph.D.

Co-Major Professor

Yasin Yilmaz, Ph.D.

Committee Member

Matthew Schabath, Ph.D.

Committee Member

Mia Naeini, Ph.D.

Committee Member

Susana Lai-Yuen, Ph.D.

Keywords

Distributed Learning, Cancer Research, Prognosis Prediction, Medical Imaging, Uncertainty Estimation

Abstract

Cancer remains one of the leading causes of mortality worldwide, and machine learning is well positioned to leverage the large volumes of routinely collected clinical and imaging data to improve patient outcomes. Realizing this potential at scale requires three capabilities that current practice does not yet deliver in combination: models that integrate the multiple data modalities clinicians use, training procedures that respect institutional data-sharing constraints, and tooling that enables sharing of de-identified data when federated approaches are not sufficient. This dissertation develops methods and infrastructure addressing each of these capabilities in the context of breast cancer and oncology more broadly.

First, a multimodal machine-learning pipeline is developed for predicting mortality risk in breast cancer patients from screening mammograms and clinical variables. The work shows that mammographic images can compensate for missing clinical biomarkers in mortality risk stratification, that the model identifies high-risk patients even within subgroups generally associated with good prognosis, and that the contralateral (unaffected) breast carries a strong, independently prognostic signal for survival of the primary cancer. This region has been historically overlooked in survival-prediction workflows.

Second, this dissertation provides a comprehensive review of federated learning for medical imaging, covering training methods, privacy-preservation techniques, and uncertainty quantification, and translates that review into practice through several real-world federatedlearning deployments at Moffitt Cancer Center as part of broader National Cancer Institute and EDRN consortia. These deployments are supported by a one-touch FL system that enables rapid conversion of existing machine-learning pipelines to a federated paradigm and the embed-thenviii federate system that allows institutions to leverage large medical foundation models without prohibitive compute and communication cost.

Third, a validated and calibrated automated de-identification pipeline for medical data containing protected health information is developed in collaboration with Impact Business Information Solutions (IBIS), with calibration that allows low-confidence cases to be flagged for human review. Taken together, these contributions advance machine learning in oncology toward multimodal models that can be trained securely across institutions and that can release clean, deidentified data when needed, ultimately helping machine learning deliver on its potential to improve cancer patient outcomes.

Share

COinS