About this Event
View mapAditya Sriram - Final Dissertation Defense - HUGEN PhD Candidate
Department of Human Genetics Doctoral Candidate, Aditya Sriram, will defend the following dissertation on “Advancing Methodological Foundations in Precision Medicine Through Novel Causal Inference, Machine Learning, and Artificial Intelligence Frameworks”
COMMITTEE CHAIR: Hyun-Jung (HJ) Park, PhD
Committee Members:
ABSTRACT:
Complex diseases arise from interactions across multiple layers of biological and clinical information, including inherited genetic components, environmental exposures, clinical contexts, and molecular regulation. This multilayered heterogeneity is a well-established concern for precision medicine, which seeks to tailor disease prevention, diagnosis, and treatment to the biological and clinical characteristics of individual patients and, within larger heterogeneous populations, more precisely defined subgroups. Machine learning models, particularly deep learning models, have become increasingly powerful for prediction in this setting. Additionally, with the advent of artificial intelligence and its growing role in healthcare and biomedical research, there is growing demand for models that extend beyond prediction toward causal inference and offer greater interpretability for researchers, clinicians, and patients. Motivated by these challenges and directions, this dissertation develops three novel deep learning-based frameworks for biologically meaningful and reliability-aware biomedical discovery in complex disease and precision medicine.
Aim 1 introduces DeepEXPOKE, a paired-reference deep-learning framework for decomposing exposure-associated disease risk into genetically driven, non-genetic, and confounding-driven components. By evaluating observed exposures against both model-X statistical knockoffs and polygenic risk score-derived matched controls, DeepEXPOKE provides a conservative framework for interpreting exposure-disease associations in large observational cohorts. Across both simulations and cohort-level analyses of sepsis, coronary heart disease, and colorectal cancer in the UK Biobank, with partial external assessment in the All of Us patient database, DeepEXPOKE showed that exposures with similar predictive value can reflect distinct underlying risk architectures that are not resolved by exposure ranking alone.
Aim 2 introduces DeepDiff-SHAP, a framework for subgroup-specific causal hypothesis generation. DeepDiff-SHAP integrates conditional SHapley Additive exPlanations (SHAP) values with a deep neural network extension of regression-based differential causal inference to identify causal relationships that change across specified subgroups. Applied to a CDC diabetes dataset, a UK Biobank sepsis cohort stratified by hypertension status, and a simulated gene expression dataset, DeepDiff-SHAP identified causal differences involving age, general health, cholesterol-related traits, alkaline phosphatase, and other clinically relevant variables, demonstrating improved sensitivity to subgroup-specific mechanisms that may be missed by linear approaches or generalized causal inference strategies.
In Aim 3, investigation of large-scale single-cell foundation models led to the introduction of scPerturbTrust, an interpretable reliability framework for in silico single-cell gene perturbation prediction experiments. Rather than explaining in silico gene perturbations based on predicted gene-expression changes alone, scPerturbTrust evaluates whether a specific gene perturbation query is likely to yield a reliable prediction and explains why. Using query-level features that capture model exposure, cellular context, and molecular regulatory properties, scPerturbTrust identified determinants of perturbation reliability rooted in both model exposure and training context as well as the biological properties of the genes involved. Validation across distinct single-cell foundation model systems further supported the presence of transferable, biologically interpretable reliability components.
Together, the three aims of this dissertation establish a unified framework for trustworthy biomedical machine learning. Through stratification of exposure-associated risk, identifying subgroup-specific causal changes in heterogeneous patient populations and data, and assessing query-level reliability for gene perturbation prediction in large foundation models, this work advances interpretable computational frameworks centered around deep learning and artificial intelligence as clinically meaningful tools for precision medicine and biological discovery.
IN-PERSON EVENT - ALL WELCOME
Please let us know if you require an accommodation in order to participate in this event. Accommodations may include live captioning, ASL interpreters, and/or captioned media and accessible documents from recorded events. At least 5 days in advance is recommended.